Topic 3.8 Notes – Operant Conditioning
How Operant Conditioning Works
Operant conditioning connects a behavior to its consequence. If the consequence strengthens the behavior, that behavior is more likely to happen again. If it weakens the behavior, it becomes less likely.
Thorndike laid the groundwork with the Law of Effect. He found that behaviors followed by satisfying outcomes get repeated, and behaviors followed by unpleasant outcomes fade.
Thorndike and Skinner
- Thorndike’s puzzle box used cats that escaped by trial and error.
- At first, behavior was random.
- Over trials, the cat learned the action that opened the box.
- Escape time decreased, showing learning.
- B. F. Skinner studied this more precisely with the Skinner box.
- Rats pressed levers.
- Pigeons pecked keys.
- These are emitted behaviors, meaning voluntary actions the organism performs.
The diagram below shows a typical Skinner box, or operant chamber, with the rat, response lever, food dispenser, signal lights, and grid floor used to study how consequences shape behavior.

Skinner box
A key AP distinction here is this:
- Operant conditioning = behavior-consequence association
- Classical conditioning = stimulus-stimulus association
Also, contingency matters. The consequence has to depend on the behavior. If food appears no matter what the rat does, that is not true operant conditioning.
One more thing students mix up a lot: a reinforcer is only a reinforcer if behavior increases. A punisher is only a punisher if behavior decreases. The label depends on the result, not the intention. Immediate, consistent consequences usually create the clearest learning.
Types of Consequences
Reinforcement and punishment
There are two questions to ask:
- Did the behavior increase or decrease?
- Was something added or removed?
| Term | What happens to behavior | What happens to stimulus |
|---|---|---|
| Positive reinforcement | Increases | Something added |
| Negative reinforcement | Increases | Something removed |
| Positive punishment | Decreases | Something added |
| Negative punishment | Decreases | Something removed |
Positive means added. Negative means removed. It does not mean good or bad.
Positive reinforcement
Something desirable is added, and behavior increases.
- Food after a lever press
- Praise after studying
Negative reinforcement
Something aversive is removed or prevented, and behavior increases.
- Escape learning ends an unpleasant stimulus already happening
- Avoidance learning prevents it from happening
Example:
- Buckling your seatbelt stops the annoying buzzer
Positive punishment
Something aversive is added, and behavior decreases.
- Touching a hot stove leads to pain, so that behavior decreases
Negative punishment
Something valued is taken away, and behavior decreases.
- Losing phone privileges after breaking a rule
Primary and secondary reinforcers
- Primary reinforcers have natural biological value
- food, water, warmth
- Secondary reinforcers get their value through learning
- money, grades, tokens, praise
Cues, Shaping, and Limits on Learning
Discrimination and generalization
A discriminative stimulus is a cue that signals when a behavior will likely be reinforced.
- If a rat gets food only when a green light is on, it learns to press then and not at other times. That is reinforcement discrimination.
- If it also presses when a similar blue-green light appears, that is reinforcement generalization.
Shaping
Shaping means reinforcing successive approximations of a target behavior.
A classic sequence looks like this, moving step by step toward the final response:
- Reinforce approaching the lever
- Reinforce getting close to the lever
- Reinforce touching the lever
- Reinforce pressing the lever

Shaping through successive approximations
This is useful when the full behavior does not happen on its own at first.
Instinctive drift
Biology can interfere with conditioning. Instinctive drift happens when trained behavior shifts toward natural, species-typical behavior.
- Brelands’ raccoon kept rubbing coins instead of depositing them, because rubbing matched natural food-handling behavior
Superstitious behavior and learned helplessness
- Superstitious behavior happens when an unrelated action gets accidentally reinforced.
- Skinner’s pigeons repeated odd movements that happened right before food appeared.
- Learned helplessness happens after uncontrollable aversive events.
- In Seligman and Maier’s dog study, dogs exposed to unavoidable shock later failed to try escaping even when escape was possible.
The contrast matters:
- Superstition = false sense of control
- Helplessness = false lack of control
Reinforcement Schedules and Their Patterns
Continuous reinforcement
Continuous reinforcement means every correct response is reinforced.
- Best for learning a new behavior quickly
- Behavior extinguishes faster once reinforcement stops
Partial reinforcement
Partial reinforcement means only some responses are reinforced.
- Learning is slower
- Behavior is more resistant to extinction
- This is the partial-reinforcement effect
The four partial schedules
- Fixed-ratio
- reinforcement after a set number of responses
- fast responding, then a pause after reward
- Variable-ratio
- reinforcement after an unpredictable number of responses
- example: slot machines
- very fast, steady responding; most resistant to extinction
- Fixed-interval
- first response after a set amount of time
- creates a scalloped pattern
- Variable-interval
- first response after changing time intervals
- moderate, steady responding
Reading schedule graphs
A cumulative record graphs time on the x-axis and cumulative number of responses on the y-axis. The steeper the slope, the faster the responding.
In the graph below, focus on the overall line shapes for each schedule.

Cumulative record patterns for reinforcement schedules
Know these patterns:
- Fixed-ratio = stair-step
- Variable-ratio = steep and steady
- Fixed-interval = scallop
- Variable-interval = moderate steady slope
Ratio schedules usually produce faster responding than interval schedules.
Key Takeaways
Operant Conditioning
Learning in which a behavior becomes more or less likely because of the consequences that follow it
Edward Thorndike
Researcher whose puzzle-box experiments showed that consequences select behavior and led to the Law of Effect
Law of Effect
Behaviors followed by reinforcing consequences become more likely, while behaviors followed by punishing consequences become less likely
B. F. Skinner
Researcher who developed the systematic experimental study of operant conditioning
Operant Chamber (Skinner Box)
An enclosure that records an animal’s responses and delivers programmed consequences for studying operant conditioning
Contingency
A relationship in which a consequence occurs because of, or is conditional on, a particular behavior
Reinforcement
A consequence that increases the future frequency or probability of the behavior it follows
Punishment
A consequence that decreases the future frequency or probability of the behavior it follows
Positive Reinforcement
Adding a stimulus after a behavior so that the behavior increases
Negative Reinforcement
Removing, reducing, or preventing an aversive stimulus after a behavior so that the behavior increases
Escape Response vs. Avoidance Response
An escape response ends an aversive event already occurring; an avoidance response prevents an expected aversive event
Positive Punishment
Adding a stimulus after a behavior so that the behavior decreases
Negative Punishment
Removing a valued stimulus after a behavior so that the behavior decreases
Primary Reinforcer
A reinforcer with unlearned value because it satisfies a biological need or innate form of comfort
Secondary Reinforcer (Conditioned Reinforcer)
A reinforcer that acquires value through learned association with other reinforcers
Discriminative Stimulus
An antecedent cue signaling that a particular response is likely to be reinforced in its presence
Reinforcement Discrimination
Learning to respond under a stimulus condition in which the behavior is reinforced but not under conditions in which it is not
Reinforcement Generalization
Performing a learned behavior in the presence of stimuli similar to the cue under which it was reinforced
Shaping
Developing a target behavior by reinforcing successive approximations that come progressively closer to it
Instinctive Drift
The tendency for a trained animal behavior to shift toward an instinctive behavior that competes with the trained response
Superstitious Behavior
Behavior repeated because it was accidentally followed by reinforcement despite lacking a genuine response-consequence contingency
Learned Helplessness
Reduced attempts to escape or exert control after learning through repeated uncontrollable aversive events that responding is ineffective
Reinforcement Schedule
The rule determining which occurrences of a behavior will receive reinforcement
Continuous Reinforcement Schedule
Reinforcement of every occurrence of the target behavior, producing rapid acquisition but relatively rapid extinction
Partial Reinforcement Schedule (Intermittent Reinforcement Schedule)
Reinforcement of only some occurrences of a behavior, usually producing slower acquisition but greater resistance to extinction
Partial-Reinforcement Effect
The greater resistance to extinction of behavior learned under partial reinforcement than under continuous reinforcement
Operant Extinction
The decline of a previously reinforced operant response after reinforcement is discontinued
Ratio vs. Interval Schedules
Ratio schedules depend on the number of responses; interval schedules reinforce the first qualifying response after time has passed
Fixed vs. Variable Schedules
Fixed schedules use predictable requirements; variable schedules use unpredictable requirements organized around an average
Fixed-Ratio Schedule
Reinforcement after a set number of responses, producing high responding with postreinforcement pauses
Variable-Ratio Schedule
Reinforcement after an unpredictable number of responses around an average, producing very high, steady, extinction-resistant responding
Fixed-Interval Schedule
Reinforcement of the first response after a set time, producing a scalloped pattern of pausing and accelerating responding
Variable-Interval Schedule
Reinforcement of the first response after unpredictable time intervals around an average, producing moderate, steady responding
Cumulative Record
A graph with time on the horizontal axis and cumulative responses on the vertical axis, whose slope shows response rate
Notes
Operant Conditioning
Learning in which a behavior becomes more or less likely because of the consequences that follow it
Edward Thorndike
Researcher whose puzzle-box experiments showed that consequences select behavior and led to the Law of Effect
Law of Effect
Behaviors followed by reinforcing consequences become more likely, while behaviors followed by punishing consequences become less likely
B. F. Skinner
Researcher who developed the systematic experimental study of operant conditioning
Operant Chamber (Skinner Box)
An enclosure that records an animal’s responses and delivers programmed consequences for studying operant conditioning
Contingency
A relationship in which a consequence occurs because of, or is conditional on, a particular behavior
Reinforcement
A consequence that increases the future frequency or probability of the behavior it follows
Punishment
A consequence that decreases the future frequency or probability of the behavior it follows
Positive Reinforcement
Adding a stimulus after a behavior so that the behavior increases
Negative Reinforcement
Removing, reducing, or preventing an aversive stimulus after a behavior so that the behavior increases
Escape Response vs. Avoidance Response
An escape response ends an aversive event already occurring; an avoidance response prevents an expected aversive event
Positive Punishment
Adding a stimulus after a behavior so that the behavior decreases
Negative Punishment
Removing a valued stimulus after a behavior so that the behavior decreases
Primary Reinforcer
A reinforcer with unlearned value because it satisfies a biological need or innate form of comfort
Secondary Reinforcer (Conditioned Reinforcer)
A reinforcer that acquires value through learned association with other reinforcers
Discriminative Stimulus
An antecedent cue signaling that a particular response is likely to be reinforced in its presence
Reinforcement Discrimination
Learning to respond under a stimulus condition in which the behavior is reinforced but not under conditions in which it is not
Reinforcement Generalization
Performing a learned behavior in the presence of stimuli similar to the cue under which it was reinforced
Shaping
Developing a target behavior by reinforcing successive approximations that come progressively closer to it
Instinctive Drift
The tendency for a trained animal behavior to shift toward an instinctive behavior that competes with the trained response
Superstitious Behavior
Behavior repeated because it was accidentally followed by reinforcement despite lacking a genuine response-consequence contingency
Learned Helplessness
Reduced attempts to escape or exert control after learning through repeated uncontrollable aversive events that responding is ineffective
Reinforcement Schedule
The rule determining which occurrences of a behavior will receive reinforcement
Continuous Reinforcement Schedule
Reinforcement of every occurrence of the target behavior, producing rapid acquisition but relatively rapid extinction
Partial Reinforcement Schedule (Intermittent Reinforcement Schedule)
Reinforcement of only some occurrences of a behavior, usually producing slower acquisition but greater resistance to extinction
Partial-Reinforcement Effect
The greater resistance to extinction of behavior learned under partial reinforcement than under continuous reinforcement
Operant Extinction
The decline of a previously reinforced operant response after reinforcement is discontinued
Ratio vs. Interval Schedules
Ratio schedules depend on the number of responses; interval schedules reinforce the first qualifying response after time has passed
Fixed vs. Variable Schedules
Fixed schedules use predictable requirements; variable schedules use unpredictable requirements organized around an average
Fixed-Ratio Schedule
Reinforcement after a set number of responses, producing high responding with postreinforcement pauses
Variable-Ratio Schedule
Reinforcement after an unpredictable number of responses around an average, producing very high, steady, extinction-resistant responding
Fixed-Interval Schedule
Reinforcement of the first response after a set time, producing a scalloped pattern of pausing and accelerating responding
Variable-Interval Schedule
Reinforcement of the first response after unpredictable time intervals around an average, producing moderate, steady responding
Cumulative Record
A graph with time on the horizontal axis and cumulative responses on the vertical axis, whose slope shows response rate