Topic 2.2 Notes – Data Compression
1. What Data Compression Is
Data compression is the process of reducing the number of bits used to store or transmit data.
When you compress a file:
- The meaning of the data stays the same for the user.
- The internal representation (the bits) changes.
- The total number of bits usually decreases.
This connects to a big idea in Unit 2: computers store everything as binary. Compression changes the binary representation, not the idea behind it.
Fewer Bits ≠ Less Information
This line confuses people, so let’s make it clear.
If data contains repeated patterns, we can replace those patterns with shorter codes. The information is still there, just represented more efficiently.
Here’s a simple example:
Original: AAAAAABBBBCC
Compressed (idea): 6A4B2C
You used fewer characters, but you could reconstruct the exact original. No information was lost because the pattern was predictable.
Why Compression Matters
Large files (images, audio, video):
- Use lots of storage space
- Take longer to download or upload
- Use more bandwidth on networks
Smaller files mean:
- Faster streaming
- Less network congestion
- Lower storage costs
On AP-style questions, if they mention “limited bandwidth” or “large media files,” your brain should immediately think about compression trade-offs.
2. What Determines How Much a File Can Be Compressed
Two factors control how small a file can get.
Redundancy in the Data
Redundancy means repeated or predictable patterns.
Examples:
- A photo with large areas of the same color (like a blue sky)
- A text file where certain words appear often
- Long runs of identical pixels in simple graphics
More repetition → more compressible
Less repetition (random-looking data) → less compressible
If data is already highly random, there’s almost nothing to “shrink.”
The Compression Algorithm Used
A compression algorithm is the method used to detect and encode patterns.
Different algorithms:
- Look for different types of patterns
- Replace them in different ways
- Produce different final file sizes
Same file + different algorithm = different compression result.
So compression effectiveness depends on:
The structure of the data × the strategy of the algorithm
AP questions often describe two algorithms and ask which produces a smaller file. The answer depends on which one better matches the data’s redundancy.
3. Types of Data Compression
There are two major categories you must clearly distinguish.
Lossless vs Lossy
| Feature | Lossless Compression | Lossy Compression |
|---|---|---|
| Reconstruction | Original can be rebuilt exactly | Only an approximation is rebuilt |
| Data removed? | No data permanently removed | Some data permanently discarded |
| Size reduction | Usually smaller reduction | Usually much greater reduction |
| Common uses | Text, programs, medical data | Photos, audio, video |
Lossless Compression
With lossless, nothing is permanently deleted.
- Removes redundancy only
- Guarantees exact reconstruction
- Essential when precision matters
Examples of when it’s chosen:
- Financial records
- Medical imaging
- Executable programs
- Databases
If even one bit changes in a program, it might not run. That’s why lossy is not acceptable there.
Lossy Compression
With lossy, some data is permanently removed.
- Reconstructed version is close, but not identical
- Often removes details humans don’t easily notice
- Achieves greater size reduction
Common uses:
- Streaming video
- Music files
- Social media images
Here’s the key insight: lossy compression works because human perception has limits. Slight color or sound differences may not be noticeable.
On test questions, if they say “maximize quality” or “exact reconstruction required,” pick lossless.
If they say “minimize download time” or “limited storage,” lossy is usually the better choice.
4. Choosing the Right Compression for a Situation
This is the skill they actually test: comparing algorithms in context.
When quality and accuracy are most important
Choose lossless.
Situations:
- Scientific measurements
- Legal documents
- Source code files
- Medical scans
The priority is exact reconstruction.
When minimizing size or transmission time is most important
Choose lossy.
Situations:
- Streaming movies over slow internet
- Sharing large photo collections
- Uploading videos to social media
The priority is smaller file size and faster transfer.
The Trade-Off
You’re always balancing:
- ✔️ Perfect accuracy
- ✔️ Smaller file size
You rarely get both at maximum levels.
If a question describes a system with strict storage limits, expect lossy.
If it describes critical data integrity, expect lossless.