AI Programming
Psychology in Tech
Exploring the intersection of psychology, programming, statistics, and data science
Discover how psychological principles enhance software development, data analysis, and user experience design
Why Psychology Matters in Technology
Psychology plays a crucial role in technology development, from understanding user behavior to designing intuitive interfaces. By applying psychological principles, developers can create more effective, user-friendly, and ethically sound technological solutions.
This guide explores how psychology intersects with programming, coding practices, statistical analysis, and data science to create better digital experiences and more meaningful insights.
Research Data Analysis
Use Python/R for statistical analysis in psychology research projects. Analyze survey data, experimental results, and behavioral patterns using statistical methods.
Thesis & Dissertation Work
Apply programming to automate data collection, clean datasets, and visualize research findings. Use version control (Git) to track your research progress.
Statistical Software
Learn SPSS, R, or Python for statistical analysis. Use these tools to run t-tests, ANOVAs, regression analyses, and other psychological research methods.
As a University Student
Research Data Analysis
Use Python/R for statistical analysis in psychology research projects. Analyze survey data, experimental results, and behavioral patterns using statistical methods.
Thesis & Dissertation Work
Apply programming to automate data collection, clean datasets, and visualize research findings. Use version control (Git) to track your research progress.
Statistical Software
Learn SPSS, R, or Python for statistical analysis. Use these tools to run t-tests, ANOVAs, regression analyses, and other psychological research methods.
Data Visualization
Create compelling visualizations of psychological data using matplotlib, ggplot2, or Tableau. Present your research findings effectively to professors and peers.
Experimental Design
Use programming to design and run online experiments, surveys, and behavioral studies. Automate data collection and analysis workflows.
Collaborative Projects
Use Git/GitHub to collaborate on group projects. Learn to manage code, share data analysis scripts, and work together on research assignments.
As an Academic/Professor
Research Publication
Use R/Python for reproducible research. Create analysis scripts that can be shared with reviewers and other researchers for transparency and verification.
Teaching & Course Materials
Develop interactive learning materials using Jupyter notebooks, R Markdown, or Shiny apps. Create engaging statistics and research methods courses.
Grant Applications
Use data science skills to strengthen grant proposals. Demonstrate expertise in statistical analysis and data management in funding applications.
Student Supervision
Guide students in using programming and statistics for their research. Review code, provide feedback on data analysis guidance, and ensure best practices.
Research Collaboration
Use version control systems to collaborate with international research teams. Share code, data analysis pipelines, and research workflows efficiently.
Open Science Practices
Promote reproducible research by sharing code, data, and analysis scripts. Use GitHub, OSF, or other platforms to make research transparent and accessible.
Coding for Beginners: Getting Started with Psychology & Programming
A comprehensive guide for psychology students and researchers new to programming
Start with the Basics
Learn fundamental programming concepts like variables, data types, and basic operations. Understanding these concepts is like learning the alphabet before writing sentences.
Choose Your First Language
For psychology students, Python or R are excellent starting points. Python is versatile and beginner-friendly, while R is specifically designed for statistical analysis.
Practice with Simple Exercises
Begin with basic exercises like calculating averages, working with lists, and simple data manipulation. These skills directly apply to psychological research.
Work with Real Data
Start analyzing actual psychological data - survey responses, experimental results, or behavioral measurements. This makes learning practical and relevant.
Learn Data Visualization
Create simple charts and graphs to visualize your data. Visualization helps you understand patterns and communicate findings effectively.
Join a Community
Connect with other psychology students learning to code. Join online forums, attend workshops, and don't be afraid to ask questions. Everyone starts as a beginner!
Here are practical projects you can start with to learn coding while working on psychology-related tasks:
Survey Data Analyzer
Create a simple program to calculate basic statistics from survey responses
# Python example
import statistics
scores = [85, 90, 78, 92, 88]
mean = statistics.mean(scores)
print(f'Average score: {mean}')Simple Experiment Simulator
Simulate a basic psychology experiment with random data generation
# Python example
import random
# Simulate coin flip experiment
flips = [random.choice(['H', 'T']) for _ in range(100)]
heads = flips.count('H')
print(f'Heads: {heads}/100')Data Entry Helper
Create a program to help organize and clean research data
# Python example
participants = []
participants.append({'id': 1, 'age': 25, 'score': 85})
participants.append({'id': 2, 'age': 30, 'score': 90})
print(participants)Basic Calculator for Statistics
Build a calculator that performs common statistical operations
# Python example
def calculate_mean(numbers):
return sum(numbers) / len(numbers)
scores = [85, 90, 78, 92]
print(f'Mean: {calculate_mean(scores)}')Simple Visualization Tool
Create basic charts from your research data
# Python example
import matplotlib.pyplot as plt
groups = ['Control', 'Treatment']
means = [25, 30]
plt.bar(groups, means)
plt.ylabel('Score')
plt.title('Experiment Results')
plt.show()Data Validator
Write a program to check if your data meets certain criteria
# Python example
def validate_age(age):
if 18 <= age <= 100:
return True
return False
print(validate_age(25)) # True
print(validate_age(15)) # FalseFree Online Courses
Coursera, edX, and Khan Academy offer free Python and R courses specifically for data analysis and psychology research
Interactive Tutorials
Try Codecademy, DataCamp, or freeCodeCamp for hands-on coding practice with immediate feedback
Psychology-Specific Resources
Look for 'R for Psychology' or 'Python for Psychologists' tutorials that combine programming with your field of study
Practice Platforms
Use platforms like Kaggle Learn, LeetCode (easy problems), or Project Euler to practice coding skills with real-world problems
Psychology in Statistics & Data Analysis
Understanding how human cognition affects statistical interpretation - both in general psychology and academic research
How statistics and programming are used in general psychological practice and research
Clinical Assessment Tools
Use statistical analysis to validate psychological tests, calculate reliability and validity coefficients, and interpret test scores for clinical diagnosis.
Behavioral Data Analysis
Analyze behavioral patterns, cognitive performance data, and psychological measurements using descriptive and inferential statistics.
Treatment Outcome Research
Use statistical methods to evaluate therapy effectiveness, compare treatment groups, and measure intervention outcomes in clinical settings.
Psychometric Analysis
Apply factor analysis, item response theory, and other advanced statistical methods to develop and refine psychological measurement instruments.
How statistics and programming are used in academic psychology research and teaching
Experimental Research
Design and analyze experiments using ANOVA, t-tests, and regression. Use R or Python to run statistical tests and interpret results for publication.
Meta-Analysis
Combine results from multiple studies using meta-analytic techniques. Use statistical software to calculate effect sizes and synthesize research findings.
Longitudinal Studies
Analyze data collected over time using mixed-effects models, growth curve analysis, and time-series analysis to understand developmental patterns.
Multivariate Analysis
Use advanced statistical methods like structural equation modeling, path analysis, and cluster analysis to understand complex psychological relationships.
Lesson 1: Descriptive Statistics in R
Learn how to calculate basic descriptive statistics (mean, standard deviation, etc.) for psychological survey data. This is the foundation of all statistical analysis.
# Step 1: Load required libraries
library(dplyr)
library(psych)
# Step 2: Load your survey data
# Make sure your CSV file has columns: anxiety_score, depression_score, stress_level
survey_data <- read.csv('psychology_survey.csv')
# Step 3: Calculate descriptive statistics
# This gives you mean, standard deviation, min, max, and more
results <- survey_data %>%
select(anxiety_score, depression_score, stress_level) %>%
describe() %>%
select(mean, sd, min, max, skew, kurtosis)
# Step 4: View your results
print(results)
# What each statistic means:
# mean = average value
# sd = standard deviation (how spread out the data is)
# min = lowest value
# max = highest value
# skew = how symmetric the data is
# kurtosis = how 'peaked' the distribution is✓ Analyze survey responses from 500+ participants✓ Compare anxiety levels across different age groups✓ Check data quality before running advanced analyses✓ Create summary tables for your research paperLesson 2: T-test in Python - Comparing Two Groups
Learn to compare two groups (e.g., control vs treatment) using an independent samples t-test. This is essential for experimental psychology research.
# Step 1: Import necessary libraries
import scipy.stats as stats
import pandas as pd
import numpy as np
# Step 2: Load your experimental data
# Your CSV should have columns: 'condition' (control/treatment) and 'score'
data = pd.read_csv('experiment_results.csv')
print(f'Total participants: {len(data)}')
# Step 3: Separate into two groups
control_group = data[data['condition'] == 'control']['score'].dropna()
treatment_group = data[data['condition'] == 'treatment']['score'].dropna()
print(f'Control group: n={len(control_group)}, mean={control_group.mean():.2f}')
print(f'Treatment group: n={len(treatment_group)}, mean={treatment_group.mean():.2f}')
# Step 4: Run independent samples t-test
# This tests if the two groups have significantly different means
t_stat, p_value = stats.ttest_ind(control_group, treatment_group)
# Step 5: Interpret results
print(f'\nT-test Results:')
print(f'T-statistic: {t_stat:.3f}')
print(f'P-value: {p_value:.4f}')
if p_value < 0.05:
print('✓ Groups are significantly different (p < 0.05)')
else:
print('✗ No significant difference between groups (p >= 0.05)')
# Calculate effect size (Cohen's d)
pooled_std = np.sqrt(((len(control_group)-1)*control_group.std()**2 +
(len(treatment_group)-1)*treatment_group.std()**2) /
(len(control_group) + len(treatment_group) - 2))
cohens_d = (treatment_group.mean() - control_group.mean()) / pooled_std
print(f'Effect size (Cohen\'s d): {cohens_d:.3f}')✓ Compare therapy effectiveness: treatment group vs control group✓ Test if a new teaching method improves test scores✓ Compare anxiety levels between men and women✓ Evaluate if an intervention changes behavior scoresLesson 3: ANOVA in R - Comparing Multiple Groups
Learn to compare three or more groups using one-way ANOVA. Perfect for experiments with multiple treatment conditions or comparing different populations.
# Step 1: Load required libraries
library(dplyr)
library(effectsize)
# Step 2: Load and prepare your data
# Your data should have: 'condition' (with 3+ groups) and 'score'
experiment_data <- read.csv('multi_group_experiment.csv')
# Check your groups
print(table(experiment_data$condition))
# Step 3: Run one-way ANOVA
# This tests: H0 = all groups have the same mean
# H1 = at least one group differs
model <- aov(score ~ condition, data = experiment_data)
# Step 4: View ANOVA results
summary(model)
# Look at the p-value:
# - If p < 0.05: At least one group is significantly different
# - If p >= 0.05: No significant differences between groups
# Step 5: If ANOVA is significant, run post-hoc tests
# This tells you WHICH specific groups differ from each other
if(summary(model)[[1]][["Pr(>F)"]][1] < 0.05) {
print('\nANOVA is significant! Running post-hoc tests...')
posthoc <- TukeyHSD(model)
print(posthoc)
# Interpret: Look for p.adj < 0.05 to see which groups differ
}
# Step 6: Calculate effect size (eta squared)
# This tells you how much variance is explained by group membership
eta_sq <- eta_squared(model)
print(paste('Effect size (η²):', round(eta_sq$Eta2, 3)))
# Effect size interpretation:
# η² < 0.01: small effect
# 0.01 < η² < 0.06: medium effect
# η² > 0.14: large effect✓ Compare three different therapy approaches (CBT, DBT, Control)✓ Test if test scores differ across four different teaching methods✓ Compare anxiety levels across three age groups (18-25, 26-35, 36+)✓ Evaluate intervention effectiveness across multiple treatment dosesLesson 4: Data Cleaning with Large Datasets (2000+ participants)
Master data cleaning techniques for large psychology studies. Learn to filter participants by age, consent status, remove missing data, and handle duplicates. Essential before any analysis!
# ============================================
# DATA CLEANING WORKFLOW FOR LARGE STUDIES
# ============================================
# Step 1: Import libraries
import pandas as pd
import numpy as np
# Step 2: Load your raw dataset
# Your CSV should have columns: participant_id, age, consent, score, condition, etc.
df = pd.read_csv('psychology_study_2000.csv')
print(f'📊 STEP 1: Original data loaded')
print(f' Total participants: {len(df)}')
print(f' Columns: {list(df.columns)}')
print()
# Step 3: Remove participants under 18
# This is CRITICAL for ethical compliance!
print('🔍 STEP 2: Filtering by age (18+)')
under_18 = len(df[df['age'] < 18])
print(f' Found {under_18} participants under 18 - removing...')
df = df[df['age'] >= 18]
print(f' Remaining participants: {len(df)}')
print()
# Step 4: Filter by consent
# Only analyze data from participants who gave informed consent
print('✅ STEP 3: Filtering by consent')
no_consent = len(df[df['consent'] != 'Yes'])
print(f' Found {no_consent} participants without consent - removing...')
df = df[df['consent'] == 'Yes']
print(f' Remaining participants: {len(df)}')
print()
# Step 5: Check for missing data
print('🔍 STEP 4: Checking for missing data')
missing_counts = df[['age', 'score', 'condition']].isnull().sum()
print(f' Missing values:')
for col, count in missing_counts.items():
if count > 0:
print(f' - {col}: {count} missing')
print()
# Step 6: Remove rows with missing critical data
print('🧹 STEP 5: Removing rows with missing critical data')
before_drop = len(df)
df = df.dropna(subset=['age', 'score', 'condition'])
after_drop = len(df)
print(f' Removed {before_drop - after_drop} rows with missing data')
print(f' Remaining participants: {len(df)}')
print()
# Step 7: Remove duplicates
print('🔍 STEP 6: Checking for duplicate participants')
duplicates = df.duplicated(subset=['participant_id']).sum()
print(f' Found {duplicates} duplicate entries')
df = df.drop_duplicates(subset=['participant_id'], keep='first')
print(f' Remaining participants: {len(df)}')
print()
# Step 8: Final data quality check
print('✅ STEP 7: Final data quality check')
print(f' Final dataset: {len(df)} participants')
print(f' Age range: {df["age"].min()} - {df["age"].max()}')
print(f' All have consent: {(df["consent"] == "Yes").all()}')
print(f' Missing data: {df[["age", "score", "condition"]].isnull().sum().sum()}')
print()
# Step 9: Save cleaned data
print('💾 STEP 8: Saving cleaned data')
df.to_csv('cleaned_data.csv', index=False)
print(f' ✓ Cleaned data saved to: cleaned_data.csv')
print(f' ✓ Ready for statistical analysis!')✓ Clean survey data from 2000+ university students✓ Prepare experimental data removing ineligible participants✓ Filter longitudinal study data by consent and age requirements✓ Prepare data for publication by ensuring ethical complianceLesson 5: Complete Workflow - ANOVA with Large Cleaned Dataset
Put it all together! Clean your data, then run ANOVA analysis on a large dataset. This is the complete workflow from raw data to statistical results.
# ============================================
# COMPLETE ANOVA ANALYSIS WORKFLOW
# ============================================
# Step 1: Import all necessary libraries
import pandas as pd
from scipy.stats import f_oneway
import numpy as np
from scipy import stats
# Step 2: Load your CLEANED data
# (Make sure you ran the data cleaning script first!)
df = pd.read_csv('cleaned_data.csv')
print(f'📊 Analyzing {len(df)} participants')
print(f' Conditions: {df["condition"].unique()}')
print()
# Step 3: Prepare groups for ANOVA
# Separate participants by their condition/group
control_group = df[df['condition'] == 'Control']['score'].dropna()
treatment_a = df[df['condition'] == 'Treatment A']['score'].dropna()
treatment_b = df[df['condition'] == 'Treatment B']['score'].dropna()
# Step 4: Check your groups before analysis
print('📈 Group Summary Statistics:')
print(f' Control Group:')
print(f' n = {len(control_group)}')
print(f' Mean = {control_group.mean():.2f}')
print(f' SD = {control_group.std():.2f}')
print()
print(f' Treatment A:')
print(f' n = {len(treatment_a)}')
print(f' Mean = {treatment_a.mean():.2f}')
print(f' SD = {treatment_a.std():.2f}')
print()
print(f' Treatment B:')
print(f' n = {len(treatment_b)}')
print(f' Mean = {treatment_b.mean():.2f}')
print(f' SD = {treatment_b.std():.2f}')
print()
# Step 5: Run one-way ANOVA
# This tests: Are the group means significantly different?
print('🔬 Running One-Way ANOVA...')
f_stat, p_value = f_oneway(control_group, treatment_a, treatment_b)
print(f'\n📊 ANOVA Results:')
print(f' F-statistic: {f_stat:.3f}')
print(f' P-value: {p_value:.6f}')
print()
# Step 6: Interpret significance
if p_value < 0.05:
print('✅ SIGNIFICANT RESULT (p < 0.05)')
print(' At least one group is significantly different from others')
print(' → You should run post-hoc tests to see which groups differ')
else:
print('❌ NOT SIGNIFICANT (p >= 0.05)')
print(' No significant differences between groups')
print()
# Step 7: Calculate effect size (eta squared)
# This tells you HOW MUCH the groups differ (not just IF they differ)
print('📏 Calculating Effect Size (η²)...')
# Calculate sum of squares between groups
ss_between = sum([len(group) * (group.mean() - df['score'].mean())**2
for group in [control_group, treatment_a, treatment_b]])
# Calculate total sum of squares
ss_total = sum((df['score'] - df['score'].mean())**2)
# Eta squared = variance explained by group membership
eta_squared = ss_between / ss_total
print(f' Effect size (η²): {eta_squared:.3f}')
# Interpret effect size
if eta_squared < 0.01:
effect_size = 'small'
elif eta_squared < 0.06:
effect_size = 'medium'
else:
effect_size = 'large'
print(f' Effect size interpretation: {effect_size} effect')
print()
# Step 8: Summary
print('=' * 50)
print('SUMMARY')
print('=' * 50)
print(f'Total participants analyzed: {len(df)}')
print(f'Groups compared: {len([control_group, treatment_a, treatment_b])}')
print(f'ANOVA p-value: {p_value:.4f}')
print(f'Effect size: {eta_squared:.3f} ({effect_size})')
if p_value < 0.05:
print('✓ Significant differences found - run post-hoc tests!')
else:
print('✗ No significant differences between groups')✓ Analyze large-scale intervention study with 2000+ participants across 3 conditions✓ Compare therapy effectiveness across multiple treatment approaches✓ Test if different teaching methods produce different learning outcomes✓ Evaluate intervention effects in longitudinal research with cleaned dataAI Tools for Psychology & Programming: Beginner's Guide
How AI can help you learn coding, statistics, and data analysis in psychology
Code Explanation
Ask AI to explain what code does, how functions work, or what error messages mean. Perfect for understanding concepts you're struggling with.
Debugging Help
When your code doesn't work, AI can help identify errors, suggest fixes, and explain why something went wrong.
Learning New Concepts
Use AI as a tutor to learn programming concepts, statistical methods, or data analysis techniques at your own pace.
Code Generation
Generate starter code for common tasks like data cleaning, statistical tests, or visualizations. Always review and understand the code!
Data Analysis Guidance
Get help choosing the right statistical test, interpreting results, or understanding which analysis fits your research question.
Documentation & Examples
AI can help you find relevant documentation, provide code examples, or explain how to use specific functions or libraries.
Start with Clear Questions
The better your question, the better the AI's answer. Be specific about what you want to learn or accomplish.
Use AI to Understand Errors
When you get an error, copy the full error message and your code, then ask AI to explain what went wrong and how to fix it.
Learn by Asking 'Why'
Don't just copy AI-generated code. Ask AI to explain WHY the code works, what each part does, and how you could modify it.
Practice with AI-Generated Examples
Ask AI to create practice problems or examples, then try to solve them yourself before asking for the solution.
Get Feedback on Your Code
Share your code with AI and ask for feedback on style, efficiency, or best practices. Learn to write better code!
Use AI for Learning Resources
Ask AI to recommend learning resources, explain concepts in different ways, or create study guides for topics you're learning.
Learn how to use AI to generate code for psychology research tasks. These examples show you exactly what to ask and what code you'll get back.
Data Filtering for Psychology Studies
Generate code to filter your dataset by age, consent status, and other criteria. Perfect for preparing data before analysis.
import pandas as pd
# Load the dataset
df = pd.read_csv('psychology_survey.csv')
# Step 1: Filter by age (18 or older)
df = df[df['age'] >= 18]
print(f'After age filter: {len(df)} participants')
# Step 2: Filter by consent (only 'Yes')
df = df[df['consent'] == 'Yes']
print(f'After consent filter: {len(df)} participants')
# Step 3: Remove rows with missing scores
df = df.dropna(subset=['score'])
print(f'Final dataset: {len(df)} participants')
# Save the cleaned data
df.to_csv('cleaned_survey.csv', index=False)Descriptive Statistics Calculation
Generate code to calculate mean, standard deviation, and other descriptive statistics for your research variables.
import pandas as pd
import numpy as np
# Load your data
df = pd.read_csv('psychology_data.csv')
# Calculate descriptive statistics for each variable
variables = ['anxiety_score', 'depression_score', 'stress_level']
for var in variables:
print(f'\n{var.upper()} Statistics:')
print(f' Mean: {df[var].mean():.2f}')
print(f' Standard Deviation: {df[var].std():.2f}')
print(f' Minimum: {df[var].min():.2f}')
print(f' Maximum: {df[var].max():.2f}')
print(f' Count: {df[var].count()}')
# Or use describe() for all at once
print('\nAll Descriptive Statistics:')
print(df[variables].describe())T-test for Two Groups
Generate code to compare two groups using an independent samples t-test - essential for experimental psychology.
import pandas as pd
from scipy import stats
import numpy as np
# Load experimental data
data = pd.read_csv('experiment_results.csv')
# Separate into two groups
control = data[data['condition'] == 'control']['score'].dropna()
treatment = data[data['condition'] == 'treatment']['score'].dropna()
# Print group information
print(f'Control: n={len(control)}, M={control.mean():.2f}, SD={control.std():.2f}')
print(f'Treatment: n={len(treatment)}, M={treatment.mean():.2f}, SD={treatment.std():.2f}')
# Perform t-test
t_stat, p_value = stats.ttest_ind(control, treatment)
print(f'\nT-test Results:')
print(f'T-statistic: {t_stat:.3f}')
print(f'P-value: {p_value:.4f}')
# Interpret results
if p_value < 0.05:
print('Significant difference (p < 0.05)')
else:
print('No significant difference (p >= 0.05)')
# Calculate Cohen's d (effect size)
pooled_std = np.sqrt(((len(control)-1)*control.std()**2 +
(len(treatment)-1)*treatment.std()**2) /
(len(control) + len(treatment) - 2))
cohens_d = (treatment.mean() - control.mean()) / pooled_std
print(f'Effect size (Cohen\'s d): {cohens_d:.3f}')ANOVA for Multiple Groups
Generate code to compare three or more groups using one-way ANOVA - perfect for multi-condition experiments.
import pandas as pd
from scipy.stats import f_oneway
import numpy as np
# Load data
df = pd.read_csv('multi_group_experiment.csv')
# Prepare groups
control = df[df['condition'] == 'Control']['score'].dropna()
treatment_a = df[df['condition'] == 'Treatment A']['score'].dropna()
treatment_b = df[df['condition'] == 'Treatment B']['score'].dropna()
# Print group statistics
print('Group Statistics:')
print(f'Control: n={len(control)}, M={control.mean():.2f}')
print(f'Treatment A: n={len(treatment_a)}, M={treatment_a.mean():.2f}')
print(f'Treatment B: n={len(treatment_b)}, M={treatment_b.mean():.2f}')
# One-way ANOVA
f_stat, p_value = f_oneway(control, treatment_a, treatment_b)
print(f'\nANOVA Results:')
print(f'F-statistic: {f_stat:.3f}')
print(f'P-value: {p_value:.4f}')
# Effect size (eta squared)
ss_between = sum([len(g) * (g.mean() - df['score'].mean())**2
for g in [control, treatment_a, treatment_b]])
ss_total = sum((df['score'] - df['score'].mean())**2)
eta_squared = ss_between / ss_total
print(f'Effect size (η²): {eta_squared:.3f}')
# Interpretation
if p_value < 0.05:
print('Significant differences found - run post-hoc tests!')
else:
print('No significant differences between groups')Data Visualization with Matplotlib
Generate code to create charts and graphs for visualizing your psychology research data.
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np
# Load data
df = pd.read_csv('experiment_data.csv')
# Calculate means for each group
groups = ['Control', 'Treatment A', 'Treatment B']
means = [
df[df['condition'] == 'Control']['score'].mean(),
df[df['condition'] == 'Treatment A']['score'].mean(),
df[df['condition'] == 'Treatment B']['score'].mean()
]
# Create bar chart
plt.figure(figsize=(10, 6))
plt.bar(groups, means, color=['#3498db', '#2ecc71', '#e74c3c'], alpha=0.7)
# Add labels and title
plt.xlabel('Condition', fontsize=12, fontweight='bold')
plt.ylabel('Mean Score', fontsize=12, fontweight='bold')
plt.title('Comparison of Mean Scores Across Conditions', fontsize=14, fontweight='bold')
# Add value labels on bars
for i, mean in enumerate(means):
plt.text(i, mean + 0.5, f'{mean:.2f}', ha='center', fontweight='bold')
# Add grid for easier reading
plt.grid(axis='y', alpha=0.3, linestyle='--')
# Adjust layout and save
plt.tight_layout()
plt.savefig('group_comparison.png', dpi=300, bbox_inches='tight')
plt.show()
print('Chart saved as group_comparison.png')Complete Data Cleaning Workflow
Generate a complete data cleaning script that handles multiple common issues in psychology datasets.
import pandas as pd
import numpy as np
# ============================================
# COMPLETE DATA CLEANING WORKFLOW
# ============================================
print('Starting data cleaning process...')
# Step 1: Load raw data
df = pd.read_csv('psychology_study_raw.csv')
print(f'✓ Loaded {len(df)} participants')
# Step 2: Remove participants under 18
df = df[df['age'] >= 18]
print(f'✓ After age filter: {len(df)} participants')
# Step 3: Filter by consent
df = df[df['consent'] == 'Yes']
print(f'✓ After consent filter: {len(df)} participants')
# Step 4: Handle missing data
# Remove rows with missing critical variables
critical_vars = ['age', 'score', 'condition', 'participant_id']
df = df.dropna(subset=critical_vars)
print(f'✓ After removing missing data: {len(df)} participants')
# Step 5: Remove duplicates based on participant ID
df = df.drop_duplicates(subset=['participant_id'], keep='first')
print(f'✓ After removing duplicates: {len(df)} participants')
# Step 6: Data quality check
print('\nData Quality Check:')
print(f' Age range: {df["age"].min()} - {df["age"].max()}')
print(f' All have consent: {(df["consent"] == "Yes").all()}')
print(f' Missing values: {df[critical_vars].isnull().sum().sum()}')
print(f' Duplicates: {df.duplicated(subset=["participant_id"]).sum()}')
# Step 7: Save cleaned data
df.to_csv('psychology_study_cleaned.csv', index=False)
print('\n✓ Cleaned data saved to: psychology_study_cleaned.csv')
print(f'\nFinal dataset: {len(df)} participants ready for analysis!')Always verify AI suggestions - don't blindly copy code without understanding it
Use AI to learn, not to do your work for you - the goal is to understand, not just get answers
Ask follow-up questions to deepen your understanding of concepts
Combine AI help with official documentation and textbooks for comprehensive learning
Practice writing code yourself first, then use AI to help when you're stuck
Remember: AI can make mistakes! Always test and verify AI-generated code before using it in important work
Psychology in Tech