⚖️ Research Ethics in Computer Science
⚖️ Research Ethics in Computer Science
“Science advances on trust. Every violation of research integrity corrodes the foundation of the entire enterprise.”
Research ethics in computer science spans a broad terrain: from plagiarism and data fabrication, to privacy violations in human-subjects research, to the societal impact of deployed AI systems. Understanding these norms is not optional — violations can end careers, trigger paper retractions, and cause real-world harm.
1. Academic Integrity Fundamentals
What Constitutes Misconduct
| Type | Definition | Examples |
|---|---|---|
| Fabrication | Inventing data or results that were never collected | Reporting benchmark scores that were not actually run |
| Falsification | Manipulating data or experiments to misrepresent results | Selectively deleting unfavorable data points; cherry-picking random seeds |
| Plagiarism | Using another’s ideas, text, or code without attribution | Copying related work text from another paper; using unreferenced code |
| Duplicate Publication | Submitting substantially identical work to multiple venues simultaneously | Submitting the same paper to ICML and NeurIPS at the same time |
| Gift Authorship | Including authors who did not meaningfully contribute | Adding an advisor’s name to a paper they had no role in |
| Ghost Authorship | Excluding contributors who should be listed as authors | Omitting a collaborator who wrote the methodology section |
[!CAUTION] These are not abstract rules. Fabrication and falsification are career-ending, and in some contexts (medical AI, safety-critical systems), they can cause direct physical harm. Retracted papers remain permanently marked in databases.
The Reproducibility Standard
A core ethical obligation in computational research is reproducibility: your results must be achievable by an independent researcher following your description. This requires:
- Releasing code and model weights used to produce results.
- Reporting all hyperparameters used in experiments.
- Disclosing negative results and failed runs, not just the best run.
- Specifying hardware, software, and dataset versions precisely.
See research/reproducibility.md for a full reproducibility framework.
2. Authorship and Credit
The CRediT Taxonomy
The Contributor Roles Taxonomy (CRediT) defines 14 distinct roles in research:
- Conceptualization
- Data Curation
- Formal Analysis
- Funding Acquisition
- Investigation
- Methodology
- Project Administration
- Resources
- Software
- Supervision
- Validation
- Visualization
- Writing – Original Draft
- Writing – Review & Editing
Most CS conferences and journals now support (and increasingly require) CRediT statements in submitted papers. Use them honestly.
Authorship Order Conventions in CS
| Convention | Field | Meaning |
|---|---|---|
| First author | Universally | Primary contributor (usually the PhD student) |
| Last author | ML/AI/Systems | Senior researcher / PI / advisor |
| Alphabetical | Theory / Math | Equal contribution; no seniority implied |
| Co-first authorship | Common in biology | Explicitly denoted with † symbol; equal primary contributors |
[!IMPORTANT] Discuss authorship and author order before the project begins, not after results are produced. Authorship disputes are among the most damaging conflicts in academic labs.
3. Human Subjects Research and IRB
Any research involving human participants — including user studies, surveys, interviews, dataset collection involving people, or analysis of personal data — is subject to ethical oversight.
When IRB Review Is Required
- User studies with human participants
- Analysis of data collected from people (e.g., scraped social media data, medical records)
- Crowdsourced labeling tasks (e.g., Amazon Mechanical Turk)
- Deployment of systems that interact with people in research contexts
IRB Process Overview
- Submit a protocol describing your study: purpose, participants, data collection, privacy protections, risks, and benefits.
- Receive approval (or exemption) before beginning data collection.
- Obtain informed consent from participants.
- Store data securely and anonymize where possible.
- Report any adverse events to the IRB immediately.
[!WARNING] Publishing results from human-subjects research conducted without IRB approval is grounds for paper retraction and professional sanctions, regardless of the research quality.
4. Data Ethics
Responsible Dataset Collection
| Concern | Guidance |
|---|---|
| Consent | Collect data only with explicit informed consent or clear public domain status |
| Privacy | Anonymize personally identifiable information (PII) before storage or release |
| Bias | Document demographic composition of datasets and known biases explicitly |
| Copyright | Respect intellectual property in training datasets (especially for generative models) |
| Provenance | Document data sources, collection methods, and preprocessing steps |
The Datasheet for Datasets
Google Research introduced Datasheets for Datasets — a standardized documentation format that captures:
- Motivation and intended use
- Composition and collection process
- Preprocessing and cleaning steps
- Distribution, maintenance, and legal considerations
Publishing a datasheet alongside your dataset has become best practice.
5. AI Ethics and Responsible Research
Dual-Use Research Concerns
Many CS research contributions have dual-use potential — they can be used for beneficial or harmful purposes. Examples:
- Face recognition → security applications OR surveillance abuse
- Language models → writing assistance OR disinformation generation
- Adversarial attacks → robustness research OR real-world attacks on deployed systems
- Network traffic analysis → intrusion detection OR mass surveillance
[!NOTE] Most major venues (NeurIPS, ICML, ICLR, FAccT) now require a Broader Impacts or Ethics Statement in submitted papers. This is your opportunity to discuss potential harms and mitigation strategies honestly — not to minimize concerns.
The ACM Code of Ethics
The ACM Code of Ethics and Professional Conduct is the foundational ethical framework for computer scientists. Core principles:
- Contribute to society and human well-being.
- Avoid harm.
- Be honest and trustworthy.
- Be fair and take action not to discriminate.
- Respect the work required to produce new ideas.
- Respect privacy.
- Honor confidentiality.
Fairness and Bias in Systems
Research involving machine learning must grapple with:
- Algorithmic bias: Systems that discriminate against protected groups due to biased training data or model design.
- Measurement bias: Metrics that do not capture the experience of underrepresented users.
- Deployment context: A model accurate in aggregate may be harmful at the individual level.
Recommended frameworks and tools:
- Fairlearn — Python toolkit for assessing and mitigating fairness issues.
- AI Fairness 360 (AIF360) — IBM’s comprehensive fairness toolkit.
- Model Cards — Standardized model documentation including fairness evaluation.
6. Open Science and Transparency
The Reproducibility Crisis
Meta-analyses across ML research have found that a significant fraction of published results cannot be reproduced by independent teams. Contributing causes:
- Undisclosed hyperparameter tuning
- Selective reporting of favorable random seeds
- Missing implementation details
- Dataset contamination (test set leakage)
Open Science Practices
| Practice | Tool / Platform |
|---|---|
| Preprint publication | arXiv, OpenReview |
| Code release | GitHub, Papers with Code |
| Data release | Hugging Face Datasets, Zenodo, OSF |
| Model release | Hugging Face Hub |
| Pre-registration | OSF Preregistration — commit to methodology before seeing results |
7. Citing and Attribution
What Must Be Cited
- Direct quotations
- Paraphrased ideas from other works
- Datasets used in experiments
- Code libraries and frameworks
- Statistical methods not invented by you
Citation Tools
- Zotero — Reference manager with browser plugin for one-click citation saving.
- Better BibTeX — Generates consistent citation keys for LaTeX.
- doi2bib — Convert DOI to BibTeX entry instantly.
- Google Scholar — Cite button generates BibTeX for any paper.
Further Reading
- ACM Code of Ethics
- Datasheets for Datasets — Gebru et al.
- Model Cards for Model Reporting — Mitchell et al.
- Stochastic Parrots — Bender et al. (on large language model risks)
- The Ethics of Artificial Intelligence — Bostrom & Cirkovic