- name
- academic-paper-review
- description
- Reviews academic papers for methodological rigor, argument structure, evidence quality, and scholarly conventions. Produces structured peer-review-style feedback with specific, actionable recommendations.
Use when the user asks for feedback on an academic paper, manuscript review, or scholarly critique of a research document.
Do NOT use for proofreading academic text (use proofreading), writing a paper from scratch (use research-paper-structure), or responding to peer review comments (use peer-review-response).
- license
- Apache-2.0
- metadata
- {"author":"foundry-skills","version":"1.0.0","tags":"editing academic-writing research","category":"writing","subcategory":"editing-refinement","depends":"","disclaimer":"none","difficulty":"advanced"}
# Academic Paper Review
## When to Use
**Use this skill when:**
- A user submits a complete or near-complete academic manuscript and asks for peer-review-style feedback on its merit, rigor, or publishability
- A user wants a structured critique of an empirical study, theoretical paper, systematic review, methods paper, or case study before submission to a journal or conference
- A user asks to evaluate the methodological soundness of a specific section -- such as the sampling strategy, statistical approach, or theoretical framework -- within the context of the full paper
- A user is preparing a thesis chapter or dissertation section for committee review and wants feedback calibrated to that specific academic milestone
- A user has received a rejection or major-revisions decision and wants an independent assessment of whether the reviewers' concerns are valid before responding
- A user is reviewing another scholar's manuscript (e.g., as a formal or informal peer reviewer) and wants help structuring or validating their own review
- A user needs a pre-submission check to anticipate likely reviewer objections before submitting to a specific journal
**Do NOT use this skill when:**
- The user needs only grammar, spelling, and punctuation correction -- use `proofreading` instead
- The user wants to write an academic paper from scratch or structure a new research document -- use `research-paper-structure` instead
- The user wants to draft replies to specific peer reviewer comments already received -- use `peer-review-response` instead
- The user needs citation formatting checked or a reference list corrected -- use `citation-reference` instead
- The user asks only for a summary or explanation of a paper they are reading, not a critique -- use `document-summarization` instead
- The user wants a literature review written for them rather than their existing literature section reviewed -- use `literature-review-synthesis` instead
- The user is asking about research design for a study not yet conducted -- this is methodology consulting, not manuscript review
---
## Process
### Step 1: Establish Review Context Before Reading a Single Line
Before analyzing the manuscript, gather the following parameters. They determine every subsequent judgment.
- **Discipline and subfield:** A psychology paper on cognitive load requires very different standards than a sociology paper on labor precarity. Natural sciences emphasize replicability and falsifiability; humanities emphasize interpretive depth and theoretical coherence; social sciences sit on a spectrum between them. Engineering papers may prioritize practical validation over theoretical generalization.
- **Paper type:** Empirical (quantitative, qualitative, mixed-methods), theoretical/conceptual, systematic review or meta-analysis, methods paper, case study, commentary or perspective, or replication study. Each type has distinct evaluative criteria -- a case study cannot be faulted for lacking a large sample if it correctly frames itself as exploratory.
- **Target venue:** A top-tier journal (Nature, Science, JAMA, APSR, Econometrica) demands higher rigor than a regional conference or disciplinary workshop. Know the difference between Q1 journals, field-specific A* conferences, and student or practitioner publications. If the user names the venue, calibrate accordingly.
- **Developmental stage:** Early draft (argument and structure focus), late draft (everything), pre-submission (anticipate hostile reviewers), post-rejection revision (validate or challenge reviewer critique), thesis/dissertation (committee-appropriate depth).
- **Author's stated goal:** "Make this publishable" is different from "tell me if the argument holds" which is different from "help me understand what the reviewers will object to." Clarify before reviewing.
- **Field-specific conventions:** Certain fields require specific sections (IMRAD structure for biomedical sciences, full reflexivity statements in qualitative sociology, extensive related-work sections in computer science), specific referencing styles (APA, MLA, Chicago, Vancouver, IEEE), and specific norms around authorship, acknowledgment, and ethics statements.
If the user has not provided this context, ask for it. A review without discipline context risks applying the wrong standards.
---
### Step 2: Read for Global Comprehension First -- Not Line-by-Line
Before writing a single evaluative sentence, read the full paper for holistic understanding. This prevents the common reviewer error of flagging local problems without understanding the global argument.
- Identify the **core claim** -- the single most important thing the paper asserts. Write it in one sentence. If you cannot, this is itself a major concern.
- Identify the **primary contribution** -- what does this paper add that does not already exist in the literature? Novelty can be empirical (new data, new phenomenon), methodological (new technique or application), theoretical (new framework or synthesis), or applied (new domain or practice implication).
- Identify the **research question or hypothesis** as explicitly stated, then compare it to what the paper actually investigates. Misalignment between stated question and actual investigation is one of the most common and most serious structural problems.
- Map the **logical chain**: gap in literature → research question → method choice → data collection → analysis → results → interpretation → contribution. Each link must follow from the previous. Note where the chain breaks.
- Note your **first-read impressions** before detailed analysis -- these often correspond to what a real reviewer would react to most strongly.
---
### Step 3: Evaluate the Abstract as a Standalone Document
The abstract is the paper's single most-read section and functions as an independent argument summary. Journals use abstracts for desk rejection decisions. Evaluate it as such.
- **Structured vs. unstructured abstract:** Many journals now require structured abstracts with labeled sections (Background, Methods, Results, Conclusions). If the venue requires this and the paper does not use it, flag it.
- **The five-element test:** A high-quality abstract contains (1) the problem or gap being addressed, (2) the specific research question or objective, (3) the method used, (4) the key quantitative or qualitative results -- not just "results are presented" but actual findings, and (5) the main implication or contribution. If any element is missing, name it.
- **Word count discipline:** Most journals cap abstracts at 150-300 words. Check for violations and for padding -- abstract word count is finite and every sentence must earn its place.
- **No undefined abbreviations or unexplained jargon:** Abstracts are read by editors and reviewers who may not be deep specialists. Undefined technical terms in the abstract are a red flag.
- **Promises vs. delivery:** The abstract must not promise findings the paper does not deliver. Check every claim in the abstract against the actual results section.
---
### Step 4: Audit the Introduction and Literature Review
The introduction and literature review establish the intellectual warrant for the study. Failures here undermine everything downstream.
- **The gap-claim test:** The introduction must demonstrate a specific, non-trivial gap in existing knowledge -- not "this topic has not been studied" (too vague) but "prior studies have examined X using method Y in context Z, but no study has examined X in context W, which matters because..." Flag introductions that claim novelty without specifying the precise gap.
- **Citation recency and depth:** In fast-moving fields (neuroscience, machine learning, epidemiology), cite primarily within the last 5 years with seminal older works. In slower-moving fields (classical history, pure mathematics), older foundational works are appropriate. A literature review citing only sources older than 10 years in an active field is a red flag.
- **Synthesis vs. annotation:** A strong literature review synthesizes sources thematically -- grouping by finding, argument, or methodology -- rather than listing papers sequentially ("Smith found X. Jones found Y. Brown found Z."). An annotated bibliography masquerading as a literature review is a common weakness in student and early-career work.
- **Self-citation bias:** If more than 30-40% of references are from the same author group, flag potential self-citation inflation. This is a known problem in certain subfields.
- **The theoretical framework:** Does the paper ground itself in an established theoretical framework, and is that choice justified? A study of organizational behavior should situate itself within relevant theory (agency theory, social capital theory, institutional theory) rather than operating atheoretically. Flag papers that apply a framework mechanically without examining its fit.
- **Research question precision:** The research question must be specific enough to be falsifiable or answerable. "What are the effects of social media?" is not a research question. "Does Instagram use frequency predict self-reported body dissatisfaction in women aged 18-25 after controlling for pre-existing eating disorder risk?" is a research question.
---
### Step 5: Conduct a Rigorous Methodology Audit
Methodology is the single most scrutinized section in empirical peer review. Evaluate it using the specific standards of the paper's methodological tradition.
**For quantitative research:**
- **Sample size and power:** Was an a priori power analysis conducted? For most behavioral research, adequate power (β = 0.80) for detecting a medium effect (Cohen's d = 0.5) requires approximately 64 participants per group. For small effects (d = 0.2), this rises to ~394 per group. Undersized studies cannot reliably detect effects; oversized studies may detect trivial effects as statistically significant.
- **Sampling strategy and representativeness:** Convenience samples (undergraduate students, MTurk workers, single organization) severely limit generalizability. The paper must acknowledge this. If the sample is convenience-based, check whether the conclusions are appropriately hedged.
- **Measurement validity:** Are the measures validated instruments or ad hoc scales? Validated scales (e.g., PHQ-9 for depression, Big Five Inventory for personality) should be cited with their psychometric properties. Novel scales must report internal consistency (Cronbach's α ≥ 0.70 is the common threshold), test-retest reliability, and ideally factor structure.
- **Statistical assumptions:** Check whether required assumptions are tested -- normality for parametric tests, homogeneity of variance for ANOVA, multicollinearity for regression (VIF < 10, ideally < 5), independence of observations. Violations that are not addressed or acknowledged are a major concern.
- **Effect sizes vs. p-values:** A paper reporting only p < 0.05 without effect sizes (Cohen's d, η², r, OR, β) is methodologically incomplete. Statistical significance without effect size tells you nothing about practical importance. Conversely, large effect sizes with p > 0.05 in underpowered studies may reflect real effects obscured by inadequate sample size.
- **Multiple comparisons correction:** Studies running multiple statistical tests without correction (Bonferroni, Benjamini-Hochberg, FDR) inflate the false-positive rate. Each test at α = 0.05 gives a 5% false-positive rate; 20 uncorrected tests yield approximately 1 false positive by chance alone.
- **Confound control:** Are plausible confounding variables identified and controlled for statistically (regression covariates, propensity score matching) or by design (randomization, matching)?
**For qualitative research:**
- **Paradigmatic consistency:** The methodology must be internally consistent. A phenomenological study should follow phenomenological analysis (IPA, Giorgi method); grounded theory requires theoretical sampling and constant comparative analysis; ethnography requires extended field presence. Applying quantitative concepts like "representativeness" to purposive qualitative samples reflects a category error.
- **Trustworthiness criteria:** Qualitative rigor is assessed through Lincoln and Guba's criteria -- credibility (member checking, prolonged engagement, peer debriefing), transferability (thick description), dependability (audit trail), and confirmability (reflexivity statement). Check whether the paper addresses these.
- **Sample sufficiency:** Qualitative sample sizes should be justified by saturation, not by convention. Thematic saturation in interview studies typically occurs between 12-20 participants for homogeneous samples; heterogeneous samples may require more. Papers reporting 5 interviews as "sufficient" without saturation discussion require scrutiny.
- **Reflexivity:** Qualitative researchers must account for their positionality -- how their identity, assumptions, and prior knowledge shaped data collection and interpretation. A qualitative paper without a reflexivity section or statement is methodologically incomplete in most social science and humanities traditions.
- **Analytic transparency:** The reader should be able to trace how raw data (interview transcripts, field notes, documents) were transformed into themes or categories. Papers that jump from "we collected data" to "we found three themes" without describing the analytic process fail basic transparency standards.
**For systematic reviews and meta-analyses:**
- **PRISMA compliance:** Systematic reviews should follow the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 checklist. A PRISMA flow diagram showing screening stages is expected.
- **Search strategy reproducibility:** The search strategy must be reported in enough detail to be replicated -- specific databases searched, search terms and Boolean operators used, date ranges, language restrictions, inclusion/exclusion criteria.
- **Risk of bias assessment:** Studies included in a meta-analysis should be assessed for quality/risk of bias using validated tools (Cochrane Risk of Bias tool for RCTs, Newcastle-Ottawa Scale for observational studies, ROBIS for systematic reviews). Failure to do this is a major methodological gap.
- **Heterogeneity:** Meta-analyses must report statistical heterogeneity (I² statistic). I² < 25% indicates low heterogeneity; 25-75% moderate; > 75% high. High heterogeneity may render a pooled effect size misleading and requires subgroup analysis or narrative synthesis.
- **Publication bias:** Check for funnel plot asymmetry, Egger's test, or sensitivity analyses addressing the file-drawer problem. A meta-analysis that does not address publication bias is incomplete.
---
### Step 6: Evaluate Results and Discussion Sections
- **Results neutrality:** The results section must present findings without interpretation. Analysis belongs in the discussion. A results section that editorializes ("This surprising finding suggests...") confounds the two sections.
- **Tables and figures:** Every table and figure must be interpretable standalone -- with title, labeled axes, units, and notes. Data presented in both tables and prose without additional insight is redundant. Figures must add something that prose or tables do not.
- **Discussion scope:** The discussion must interpret results in light of the research question (not restate results), situate findings in the existing literature (how do they confirm, challenge, or extend prior work?), and acknowledge limitations honestly and specifically.
- **Limitations section quality:** Vague limitations ("larger samples are needed") are unhelpful. Specific limitations name the exact threat to validity, explain its magnitude and direction (does it inflate or deflate the effect?), and explain why it could not be fully addressed. A strong limitations section actually strengthens reviewer confidence.
- **Contribution clarity:** The paper must explicitly state what it contributes and why that contribution matters. This should appear in the discussion/conclusion and be traceable back to the gap identified in the introduction. If the stated contribution does not match the gap, flag the misalignment.
- **Overgeneralization:** Claims must be scoped to the sample and context studied. A study of 200 university students in the United States cannot conclude "humans tend to..." or "people generally..." without extraordinary evidence. Every overclaim of generalizability is a major concern.
---
### Step 7: Assess Ethics, Transparency, and Scholarly Integrity
These elements are increasingly required by journals and are legitimate review concerns.
- **Ethics approval:** Empirical research involving human participants should report IRB (Institutional Review Board) or equivalent ethics board approval with an approval number or statement. Animal research requires IACUC or equivalent. The absence of an ethics statement is a publishability barrier at most journals.
- **Informed consent:** Human subjects research should confirm that informed consent was obtained. If data are archival, secondary, or publicly available, this should be stated.
- **Data availability:** Many journals now require a data availability statement. Does the paper indicate where the data can be accessed, or justify why it cannot be shared (confidentiality, proprietary)?
- **Pre-registration:** In psychology, medicine, and increasingly other fields, pre-registration of hypotheses and analysis plans (via OSF, ClinicalTrials.gov, AsPredicted) is valued because it distinguishes confirmatory from exploratory analysis. If the paper is not pre-registered and makes confirmatory claims, flag the absence.
- **Conflict of interest:** Does the paper disclose funding sources and potential conflicts of interest? Industry-funded studies showing positive results for the funder's product warrant scrutiny.
- **Acknowledgments:** Are contributors who do not meet authorship criteria appropriately acknowledged?
---
### Step 8: Synthesize and Structure the Review
After completing the analysis, structure feedback in a formal peer review format.
- Write the **summary first** -- this demonstrates to the author that you understood their work before critiquing it. A summary that misrepresents the paper is a reviewer failure.
- Separate **major concerns** (the paper cannot be accepted as-is without addressing these) from **minor concerns** (should be addressed but would not alone preclude acceptance). This distinction matters enormously to authors and editors.
- Calibrate **recommendation language** precisely. Most journals use: Reject, Major Revision, Minor Revision, Accept with Revisions, Accept. "Major Revision" typically means the paper could be published if the authors address the identified concerns to the reviewer's satisfaction at re-review. "Reject" typically means the fundamental design, dataset, or framing cannot be fixed through revision.
- Every concern must contain: (1) the specific location (section, paragraph, or page), (2) the precise problem, (3) why it matters (impact on validity, contribution, or interpretability), and (4) a concrete recommendation.
- Frame feedback constructively. Peer review exists to improve scholarship, not to gate-keep based on personal preference. Hostile or dismissive tone is not more rigorous -- it is less useful.
---
## Output Format
```
## Academic Paper Review
**Paper Title:** [Full title as given]
**Discipline / Subfield:** [e.g., Organizational Behavior / Industrial-Organizational Psychology]
**Paper Type:** [Empirical -- Quantitative Survey / Theoretical / Systematic Review / etc.]
**Target Venue (if known):** [Journal name, conference, thesis committee]
**Review Stage:** [Pre-submission / Post-rejection revision / Thesis draft / etc.]
**Recommendation:** [Reject | Major Revision | Minor Revision | Accept with Revisions | Accept]
---
### Summary
[2-4 sentences stating: what the paper investigates, what method it uses, what it finds, and what it claims to contribute. This must demonstrate genuine understanding of the paper's purpose before any critique.]
---
### Core Contribution Assessment
[1-2 sentences assessing whether the stated contribution is real, novel, and adequately scoped. State explicitly whether the paper succeeds in filling the gap it identifies.]
---
### Strengths
1. [Specific strength, with reference to section/element -- e.g., "The cross-organizational sample design (Section 3.1) substantially improves generalizability relative to prior single-site studies."]
2. [Specific strength]
3. [Specific strength]
4. [Optional: additional strength]
5. [Optional: additional strength]
---
### Major Concerns
*(Issues that must be resolved before the paper can be accepted)*
**Concern 1: [Descriptive Title]** -- [Section/Page reference]
- **Issue:** [Precise description of the problem]
- **Why it matters:** [How this affects validity, contribution, or interpretability]
- **Recommendation:** [Specific, actionable suggestion -- not just "improve this" but how]
**Concern 2: [Descriptive Title]** -- [Section/Page reference]
- **Issue:**
- **Why it matters:**
- **Recommendation:**
[Continue for all major concerns -- typically 2-6 for a major-revision recommendation]
---
### Minor Concerns
*(Issues that should be addressed and would strengthen the paper, but do not alone prevent acceptance)*
1. [Section reference] -- [Issue] -- [Suggested fix]
2. [Section reference] -- [Issue] -- [Suggested fix]
3. [Section reference] -- [Issue] -- [Suggested fix]
[Continue as needed]
---
### Section-by-Section Notes
| Section | Assessment | Specific Comment |
|---|---|---|
| Title | [Strong / Adequate / Needs revision] | [Comment] |
| Abstract | [Strong / Adequate / Needs revision] | [Comment] |
| Introduction | [Strong / Adequate / Needs revision] | [Comment] |
| Literature Review | [Strong / Adequate / Needs revision] | [Comment] |
| Methodology | [Strong / Adequate / Needs revision] | [Comment] |
| Results | [Strong / Adequate / Needs revision] | [Comment] |
| Discussion | [Strong / Adequate / Needs revision] | [Comment] |
| Conclusion | [Strong / Adequate / Needs revision] | [Comment] |
| References | [Strong / Adequate / Needs revision] | [Comment] |
| Ethics / Transparency | [Present / Absent / Incomplete] | [Comment] |
---
### Methodological Scorecard
*(For empirical papers only)*
| Element | Status | Notes |
|---|---|---|
| Sample size / power justification | [Adequate / Inadequate / Not reported] | |
| Sampling strategy | [Probability / Convenience / Purposive -- appropriate for design?] | |
| Measurement validity | [Validated instruments / Novel scales with psychometrics / Ad hoc] | |
| Statistical assumptions tested | [Yes / Partial / No] | |
| Effect sizes reported | [Yes / No] | |
| Confound control | [Yes / Partial / No] | |
| Multiple comparisons correction | [Applied / Not applicable / Missing] | |
| Ethics approval | [Present / Absent / Not required] | |
| Data availability statement | [Present / Absent] | |
---
### Overall Assessment
[3-5 sentences. State: (a) the paper's core strength and why the topic matters, (b) the most critical issues that must be resolved, (c) the path to acceptance or publication-readiness, and (d) a candid assessment of whether the paper is close or far from publishable in the stated venue.]
```
---
## Rules
1. **Never issue a vague concern.** Every concern must identify the specific section (e.g., "Section 3.2, paragraph 4"), the precise problem (e.g., "Cronbach's α is reported for the overall scale but not for the three subscales"), and a concrete fix (e.g., "Report α for each subscale separately, as subscale-level reliability determines whether they can be used as distinct predictors"). Vague comments like "the methodology needs more detail" are not acceptable.
2. **Never apply cross-disciplinary standards inappropriately.** Do not require a qualitative ethnography to have a large random sample, do not require a historical analysis to have a hypothesis, and do not require a theoretical paper to have empirical data. Evaluate the paper against the standards of its own methodological tradition, not against the template of a different one.
3. **Always write the summary before any evaluative statement.** The summary demonstrates that you have understood the paper's purpose and argument. A critique without demonstrated comprehension is not credible peer review -- it is uninformed judgment.
4. **Always separate major from minor concerns.** Authors and editors need to know what is a deal-breaker versus what is a suggestion. Conflating them in a single list forces authors to guess at priorities, which is an imposition of the reviewer's failure on the author's time.
5. **Never recommend rejection without clearly naming the specific, irremediable flaw.** Rejection is appropriate when the fundamental design, dataset, or core argument cannot be salvaged through revision. If the problems are addressable through revision, the appropriate recommendation is Major Revision, not Reject. Misusing the Reject recommendation suppresses legitimate research.
6. **Always distinguish between statistical significance and practical significance.** A statistically significant finding (p < 0.05) may have a trivially small effect size (d = 0.05, r = 0.02). A paper that treats p < 0.05 as evidence of a meaningful effect without reporting effect sizes is methodologically incomplete, and this must be flagged.
Ver en GitHub