Topic 4.1 Notes – Ethical and Social Issues Around Data Collection
1. Privacy Risks in Data Collection and Storage
Any time a program collects personal data such as names, emails, locations, or usage history, it creates risk. Even if your algorithm is correct, the way you handle data can still harm users.
Here’s the basic flow from user input to potential risk:
Data collection and storage risk flow
Once data is stored in a database, several risks exist.
Unauthorized access
Someone who shouldn’t see the data gets access.
- Hackers breaking into a system
- Weak security practices (easy passwords, unprotected storage)
- Employees accessing data they don’t need
If data is stored, it can be exposed. No storage means no breach.
Over-collection of data
Collecting more information than the program actually needs.
Example: A homework app that asks for a student’s home address when it only needs a username.
“Maybe we’ll use it later” is not a valid reason.
Secondary use without consent
Using data for a different purpose than the user agreed to.
- Collect email for login → later use it for advertising
- Share user data with a third party without telling them
Even if the code runs perfectly, this violates trust.
Long-term storage risk
The longer data is stored, the greater the chance something goes wrong.
Old data that is no longer needed should be deleted.
Responsible Programming Practices
Ethical programmers:
- Collect only what is necessary
- Clearly explain what data is collected and why
- Require user consent
- Allow users to view, update, or delete their data
- Protect stored data using secure practices
On multiple choice, the most ethical answer almost always minimizes collection and prioritizes consent and transparency.
2. Data Quality and Why It Matters
A program is only as good as its data.
Even a perfectly written algorithm produces bad results if the dataset is flawed.
Before using a dataset, ask:
- How was this data collected?
- Who is represented?
- Who is missing?
- Is it accurate and up to date?
Poor-quality data can cause:
- Incorrect results
- Inefficient processing
- Misleading conclusions
- Unfair outcomes
Here is the key idea:

Correct algorithm + biased or incomplete data → unfair or incorrect output
The algorithm may be logically correct, but the output can still be wrong or unfair because of the input data.
3. Algorithmic Bias
Algorithmic bias means systematic and repeated errors that unfairly affect a specific group.
This usually comes from the dataset, not from obvious coding mistakes.
How bias happens
- Unrepresentative data
If certain groups are underrepresented, predictions will be less accurate for them. - Historical bias
If past data reflects unfair systems, the program may reinforce those patterns. - Design assumptions
Programmers may unintentionally build in assumptions that disadvantage some users.
For example, imagine a training dataset that looks like this:

Imbalanced training dataset, 80% Group A and 20% Group B
If 80% of training data comes from one group, the algorithm will likely perform better for that group.
Why this matters
Biased algorithms can affect:
- Hiring decisions
- Loan approvals
- Content recommendations
- Academic evaluations
The code may appear correct, but still produce unfair outcomes.
Reducing bias
- Examine how data was collected
- Test algorithms across diverse groups
- Avoid overgeneralizing from limited samples
A common exam trap is choosing a dataset that is large but not representative of the population you care about.
4. Choosing the Right Dataset for a Problem
A dataset must match the question you’re trying to answer.
Data collected for one purpose may not work for another.
Ask:
- Does the dataset directly relate to the question?
- Is it complete, or are values missing?
- Is it accurate?
- Is it representative of the target population?
For example, data from one school does not represent all students nationally.
You also cannot reliably draw conclusions beyond what the dataset supports. That’s extrapolation beyond evidence.
Common mistakes:
- Assuming correlation implies causation
- Using a dataset from one group to make claims about another
- Ignoring missing or inaccurate entries
On the exam, when asked which dataset is appropriate, pick the one that best matches the population and question being studied.
5. How Ethics Shows Up on the AP Exam
You’ll usually see this topic in scenario-based questions.
Multiple Choice
- Identify privacy violations
- Choose the most ethical data practice
- Recognize biased or inappropriate datasets
Free Response
- Design classes that check for user consent
- Limit stored personal data
- Explain why a dataset is inappropriate
- Describe how bias could affect results
The strongest answers always protect privacy, minimize unnecessary collection, acknowledge bias, and avoid unsupported conclusions.