8m left·0%
Reading Time: 8 min
Last Updated: September 10, 2026
Main Ideas: 4
Reading Time: 8 min
Last Updated: September 10, 2026
Main Ideas: 4

Topic 3.8 Notes – Operant Conditioning

Verified for 2027 AP® Psychology Exam
Read aloud
Operant conditioning is learning through consequences. A behavior happens, then what follows changes how likely that behavior is to happen again. In this topic, you need to know how reinforcement and punishment work, how behavior gets shaped, and how different reinforcement schedules create different response patterns.

How Operant Conditioning Works

Operant conditioning connects a behavior to its consequence. If the consequence strengthens the behavior, that behavior is more likely to happen again. If it weakens the behavior, it becomes less likely.

Thorndike laid the groundwork with the Law of Effect. He found that behaviors followed by satisfying outcomes get repeated, and behaviors followed by unpleasant outcomes fade.

Thorndike and Skinner

  • Thorndike’s puzzle box used cats that escaped by trial and error.
    • At first, behavior was random.
    • Over trials, the cat learned the action that opened the box.
    • Escape time decreased, showing learning.
  • B. F. Skinner studied this more precisely with the Skinner box.
    • Rats pressed levers.
    • Pigeons pecked keys.
    • These are emitted behaviors, meaning voluntary actions the organism performs.

The diagram below shows a typical Skinner box, or operant chamber, with the rat, response lever, food dispenser, signal lights, and grid floor used to study how consequences shape behavior.

Study guide illustration

Skinner box

A key AP distinction here is this:

  • Operant conditioning = behavior-consequence association
  • Classical conditioning = stimulus-stimulus association

Also, contingency matters. The consequence has to depend on the behavior. If food appears no matter what the rat does, that is not true operant conditioning.

One more thing students mix up a lot: a reinforcer is only a reinforcer if behavior increases. A punisher is only a punisher if behavior decreases. The label depends on the result, not the intention. Immediate, consistent consequences usually create the clearest learning.

Types of Consequences

Reinforcement and punishment

There are two questions to ask:

  1. Did the behavior increase or decrease?
  2. Was something added or removed?
TermWhat happens to behaviorWhat happens to stimulus
Positive reinforcementIncreasesSomething added
Negative reinforcementIncreasesSomething removed
Positive punishmentDecreasesSomething added
Negative punishmentDecreasesSomething removed

Positive means added. Negative means removed. It does not mean good or bad.

Positive reinforcement

Something desirable is added, and behavior increases.

  • Food after a lever press
  • Praise after studying

Negative reinforcement

Something aversive is removed or prevented, and behavior increases.

  • Escape learning ends an unpleasant stimulus already happening
  • Avoidance learning prevents it from happening

Example:

  • Buckling your seatbelt stops the annoying buzzer

Positive punishment

Something aversive is added, and behavior decreases.

  • Touching a hot stove leads to pain, so that behavior decreases

Negative punishment

Something valued is taken away, and behavior decreases.

  • Losing phone privileges after breaking a rule

Primary and secondary reinforcers

  • Primary reinforcers have natural biological value
    • food, water, warmth
  • Secondary reinforcers get their value through learning
    • money, grades, tokens, praise

Cues, Shaping, and Limits on Learning

Discrimination and generalization

A discriminative stimulus is a cue that signals when a behavior will likely be reinforced.

  • If a rat gets food only when a green light is on, it learns to press then and not at other times. That is reinforcement discrimination.
  • If it also presses when a similar blue-green light appears, that is reinforcement generalization.

Shaping

Shaping means reinforcing successive approximations of a target behavior.

A classic sequence looks like this, moving step by step toward the final response:

  1. Reinforce approaching the lever
  2. Reinforce getting close to the lever
  3. Reinforce touching the lever
  4. Reinforce pressing the lever

Shaping through successive approximations

This is useful when the full behavior does not happen on its own at first.

Instinctive drift

Biology can interfere with conditioning. Instinctive drift happens when trained behavior shifts toward natural, species-typical behavior.

  • Brelands’ raccoon kept rubbing coins instead of depositing them, because rubbing matched natural food-handling behavior

Superstitious behavior and learned helplessness

  • Superstitious behavior happens when an unrelated action gets accidentally reinforced.
    • Skinner’s pigeons repeated odd movements that happened right before food appeared.
  • Learned helplessness happens after uncontrollable aversive events.
    • In Seligman and Maier’s dog study, dogs exposed to unavoidable shock later failed to try escaping even when escape was possible.

The contrast matters:

  • Superstition = false sense of control
  • Helplessness = false lack of control

Reinforcement Schedules and Their Patterns

Continuous reinforcement

Continuous reinforcement means every correct response is reinforced.

  • Best for learning a new behavior quickly
  • Behavior extinguishes faster once reinforcement stops

Partial reinforcement

Partial reinforcement means only some responses are reinforced.

  • Learning is slower
  • Behavior is more resistant to extinction
  • This is the partial-reinforcement effect

The four partial schedules

  • Fixed-ratio
    • reinforcement after a set number of responses
    • fast responding, then a pause after reward
  • Variable-ratio
    • reinforcement after an unpredictable number of responses
    • example: slot machines
    • very fast, steady responding; most resistant to extinction
  • Fixed-interval
    • first response after a set amount of time
    • creates a scalloped pattern
  • Variable-interval
    • first response after changing time intervals
    • moderate, steady responding

Reading schedule graphs

A cumulative record graphs time on the x-axis and cumulative number of responses on the y-axis. The steeper the slope, the faster the responding.

In the graph below, focus on the overall line shapes for each schedule.

Study guide illustration

Cumulative record patterns for reinforcement schedules

Know these patterns:

  • Fixed-ratio = stair-step
  • Variable-ratio = steep and steady
  • Fixed-interval = scallop
  • Variable-interval = moderate steady slope

Ratio schedules usually produce faster responding than interval schedules.

Key Takeaways

A consequence only counts as reinforcement or punishment if the behavior actually changes in the expected direction.
Negative reinforcement increases behavior by removing something aversive, so it is not the same thing as punishment.
Operant conditioning depends on contingency, which means the consequence must be tied to the behavior.
Shaping works by reinforcing closer and closer versions of the target behavior, not by waiting for the full behavior to appear.
Instinctive drift shows that conditioning works within biological limits.
Superstitious behavior and learned helplessness are opposite mistakes about control.
Continuous reinforcement builds behavior quickly, but variable-ratio schedules make behavior the hardest to extinguish.
On cumulative records, fixed-interval produces a scallop and variable-ratio produces the steepest steady slope.

AP® is a trademark registered by the College Board, which is not affiliated with, and does not endorse this website.

Notes

1 credit used · 5/5 remaining