用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills --skill literature-gap-finder命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Route empirical-research requests through the Auto-Empirical Research Skills catalog when this whole repository is installed as one skill in Codex, CodeBuddy, Claude Code, or another IDE. Use to choose and load the right vendored AERS skill for causal inference, econometrics, replication, data acquisition, manuscript writing, peer review and referee responses, citation checking, de-AIGC editing, or full empirical-paper workflows without reading the entire repository at once.
中英双语学术降 AIGC / bilingual academic de-AIGC skill. Removes AI-generated writing signatures from empirical papers in economics, management, and the social sciences — in both English and Chinese. Covers Turnitin AI, GPTZero, Originality.ai on the English side and 知网 AMLC, 万方, 维普 on the Chinese side. Uses a six-step loop (intake → audit → claim-evidence check → differentiated rewrite → five-dimension self-score → cold-reader recheck) with two pattern libraries (22 English + 17 Chinese patterns), section-by-section strategies for empirical papers, and hard protections that keep every number, coefficient, and citation intact.
Use when a research task needs reproducible Kaggle discovery, metadata inspection, bounded public-data downloads, competition or kernel discovery, model discovery, or an explicitly approved Kaggle write/delete operation through the official CLI.
基于 SOC 职业分类
正在显示 SKILL.md
| name | literature-gap-finder |
| description | Method×Setting matrices and systematic gap identification |
Systematic framework for identifying research opportunities in statistical methodology
Use this skill when: positioning research contributions, finding gaps in methodology literature, identifying unexplored combinations of methods and settings, building literature reviews, or deciding on research directions.
A publishable gap must be:
| Gap Type | Description | Example |
|---|---|---|
| Method Gap | No method exists for setting | No mediation analysis for network data |
| Theory Gap | Method exists but lacks theory | Bootstrap for mediation lacks consistency proof |
| Efficiency Gap | Methods exist but are inefficient | Doubly robust mediation more efficient |
| Robustness Gap | Methods fail under violations | Mediation under measurement error |
| Computational Gap | Existing methods don't scale | Mediation with high-dimensional confounders |
| Extension Gap | Existing method needs generalization | Binary → continuous mediator |
The method-setting matrix is the core tool for finding research gaps systematically:
# Build a method-setting matrix programmatically
create_gap_matrix <- function() {
methods <- c("Regression", "Weighting/IPW", "DR/AIPW", "TMLE", "ML-based")
settings <- c("Binary treatment", "Continuous treatment",
"Time-varying", "Clustered", "High-dimensional",
"Measurement error", "Missing data", "Network")
matrix_data <- expand.grid(method = methods, setting = settings)
matrix_data$status <- "unknown" # To be filled: "developed", "partial", "gap"
matrix_data$priority <- NA
matrix_data$references <- ""
matrix_data
visualize_gaps gap_matrix
libraryggplot2
ggplotgap_matrix aesx method y setting fill status
geom_tilecolor
scale_fill_manualvalues
theme_minimal
labstitle
x y
themeaxis.text.x element_textangle hjust
Before claiming a gap, verify systematically:
| Step | Action | Tools |
|---|---|---|
| 1 | Search major databases | Google Scholar, Web of Science, Scopus |
| 2 | Search preprint servers | arXiv, bioRxiv, SSRN |
| 3 | Search R packages | CRAN, GitHub, R-universe |
| 4 | Check conference proceedings | ICML, NeurIPS, JSM, ENAR |
| 5 | Search dissertations | ProQuest, university repositories |
| 6 | Email domain experts | 2-3 experts for confirmation |
# Systematic verification checklist
verify_gap <- function(topic, keywords) {
checklist <- list(
databases_searched = c("google_scholar", "web_of_science", "pubmed", "scopus"),
search_terms = keywords,
date_range = paste(Sys.Date() - 365*5, "to", Sys.Date()),
results = list(
papers_found = 0,
closest_related = c(),
why_not_the_same = ""
),
expert_consultation = list
experts_contacted
responses
verification_status
checklist
document_verification gap_description search_log
cat
cat gap_description
cat Sys.Date
cat
db search_logdatabases_searched
cat db
cat pastesearch_logsearch_terms collapse
cat search_logverification_status
| Criterion | Weight | Score 1-5 |
|---|---|---|
| Impact (how many benefit?) | 0.25 | ___ |
| Novelty (how new?) | 0.20 | ___ |
| Tractability (can we solve it?) | 0.20 | ___ |
| Timeliness (is it hot now?) | 0.15 | ___ |
| Fit (matches our expertise?) | 0.10 | ___ |
| Publication potential | 0.10 | ___ |
Priority Score = Σ(weight × score)
# Priority scoring function
score_research_gap <- function(
impact, # 1-5: How many researchers would benefit
novelty, # 1-5: How new/original is this
tractability, # 1-5: How likely can we solve it
timeliness, # 1-5: Is this currently hot
fit, # 1-5: Matches our expertise
publication # 1-5: Publication potential
) {
weights <- c(0.25, 0.20, 0.20, 0.15, 0.10, 0.10)
scores <- c(impact, novelty, tractability, timeliness, fit, publication)
priority <- sum(weights * scores)
list(
priority_score = priority,
interpretation = case_when(
priority
priority
priority
breakdown data.frame
criterion
weight weights
score scores
weighted weights scores
rank_gaps gaps_list
scores sapplygaps_list g gpriority_score
orderscores decreasing
Systematically map methods against settings to find gaps:
METHODS
│ Regression │ Weighting │ DR/TMLE │ ML-based │
──────────┼────────────┼───────────┼─────────┼──────────│
Binary A │ ✓ │ ✓ │ ✓ │ ✓ │
Continuous│ ✓ │ ? │ ✓ │ ? │
SETTINGS ├────────────┼───────────┼─────────┼──────────│
Time-vary │ ? │ ✓ │ ✓ │ ✗ │
Clustered │ ✓ │ ? │ ? │ ✗ │
High-dim │ ✗ │ ✗ │ ? │ ✓ │
✓ = Well-developed ? = Partial/emerging ✗ = Gap
Step 1: Identify Dimensions
For mediation analysis:
| Dimension | Variations |
|---|---|
| Treatment | Binary, continuous, multi-level, time-varying |
| Mediator | Single, multiple, high-dimensional, latent |
| Outcome | Continuous, binary, count, survival, longitudinal |
| Confounding | Measured, unmeasured, time-varying |
| Structure | Single mediator, parallel, sequential, moderated |
| Data | Cross-sectional, longitudinal, clustered, network |
| Assumptions | Standard, relaxed positivity, measurement error |
Step 2: List Methods
| Method Family | Specific Methods |
|---|---|
| Regression | Baron-Kenny, product of coefficients, difference |
| Weighting | IPW, MSM, sequential g-estimation |
| Doubly Robust | AIPW, TMLE, cross-fitted |
| Semiparametric | Influence function-based |
| Bayesian | MCMC, variational |
| Machine Learning | Causal forests, DML, neural |
| Bounds | Partial identification, sensitivity |
Step 3: Fill and Analyze
Mark each cell:
│ Product │ Weighting │ DR │ Bounds │
─────────────────────────┼─────────┼───────────┼────┼────────│
2 mediators, linear │ ✓ │ ✓ │ ✓ │ ? │
2 mediators, nonlinear │ ? │ ✓ │ ? │ ✗ │
3+ mediators, linear │ ? │ ? │ ✗ │ ✗ │
3+ mediators, nonlinear │ ✗ │ ? │ ✗ │ ✗ │
With measurement error │ ✗ │ ✗ │ ✗ │ ✗ │
With unmeasured conf. │ ✗ │ ✗ │ ✗ │ ? │
Gaps identified:
Map how assumptions have been relaxed over time:
Standard Mediation (Baron-Kenny 1986)
│
┌─────────────────┼─────────────────┐
↓ ↓ ↓
No unmeasured Linearity No interaction
confounding assumed assumed
│ │ │
↓ ↓ ↓
┌───────┴───────┐ Nonparametric VanderWeele
↓ ↓ (Imai 2010) 4-way decomp
Sensitivity Bounds │
(Imai 2010) (partial ID) ↓
│ │ Multiple mediators?
↓ ↓ Longitudinal?
E-value Sharp bounds? Measurement error?
(Ding 2016) │ │
│ ↓ ↓
↓ [YOUR GAP?] [YOUR GAP?]
[YOUR GAP?]
Step 1: Identify Original Assumptions
For a classic method, list ALL assumptions:
Step 2: Trace Relaxation History
For each assumption, find papers that:
Step 3: Find Unexplored Branches
Look for:
Positivity: P(A=a|X) > ε > 0 for all a, x
│
┌───────────────┼───────────────┐
↓ ↓ ↓
Near-violation Practical Structural
positivity violations
│ │ │
↓ ↓ ↓
Trimming Overlap Extrapolation
weights assessment methods
│ │ │
↓ ↓ ↓
Truncation? Diagnostics? Bounds under
violations?
Backward: From recent key paper, trace citations:
Forward: Using Google Scholar "Cited by":
For any topic, identify:
| Category | Description | How to Find |
|---|---|---|
| Foundational | Original method papers | Most-cited, oldest |
| Textbook | Comprehensive treatments | Citations across subfields |
| Recent reviews | State-of-the-art summaries | "Review" in title, last 5 years |
| Frontier | Latest developments | Top journals, last 2 years |
| Your competition | Groups working on same gap | Recent similar titles |
1986: Baron & Kenny [foundations]
│
├──→ 1990s: SEM extensions
│
├──→ 2004: Robins & Greenland [causal foundations]
│ │
│ ├──→ 2010: Imai et al. [sensitivity]
│ │
│ ├──→ 2010: VanderWeele [4-way]
│ │ │
│ │ └──→ 2015: Book [comprehensive]
│ │
│ └──→ 2014: Tchetgen [semiparametric]
│
└──→ 2020s: ML integration [frontier]
Before claiming a gap, verify:
When you identify a gap:
## Gap: [Brief Title]
### Setting
[Precise description of the setting where the gap exists]
### Current State
- **What exists**: [Methods that partially address this]
- **What works**: [Aspects of the problem already solved]
- **What fails**: [Where current methods break down]
### The Gap
- **Precise statement**: [What is missing]
- **Why it matters**: [Who needs this, for what applications]
- **Why it's hard**: [Technical challenges]
### Evidence of Gap
- [ ] Literature search documented
- [ ] No existing solution found
- [ ] Experts consulted (optional)
### Potential Approaches
1. [Approach 1]: [Brief description]
- Pros: [Advantages]
- Cons: [Challenges]
2. [Approach 2]: [Brief description]
- Pros: [Advantages]
- Cons: [Challenges]
### Related Work
- [Paper 1]: [How it relates, why it doesn't solve gap]
- [Paper 2]: [How it relates, why it doesn't solve gap]
### Contribution Positioning
"While [existing work] addresses [related problem], no method currently
handles [specific gap]. We propose [approach] which provides [properties]."
Gap template: "[Method] assumes [simple structure], but in [application] data has [complex structure]"
Examples:
Gap template: "[Method] requires [assumption], which is violated when [situation]"
Examples:
Gap template: "When [complication], standard estimands [NDE/NIE] are not well-defined or interpretable"
Examples:
Gap template: "Efficient methods require [strong assumptions], while robust methods are inefficient"
Examples:
Gap template: "Theoretically valid approach exists but [computational limitation]"
Examples:
Strong positioning formula:
"Although [Author Year] developed [method] for [setting], their approach [limitation]. In contrast, our method [advantage] while maintaining [property]. Specifically, we contribute: (1) [theoretical contribution], (2) [methodological contribution], (3) [practical contribution]."
| Position | When to Use | Example Language |
|---|---|---|
| Extension | Build on existing | "We extend [method] to [new setting]" |
| Synthesis | Combine approaches | "We unify [method A] and [method B]" |
| Alternative | Different approach | "We propose an alternative that [advantage]" |
| Correction | Fix limitation | "We address the limitation of [method]" |
| Generalization | Broader framework | "We develop a general framework that includes [special cases]" |
| Dimension | Competitor 1 | Competitor 2 | Our Method |
|---|---|---|---|
| Setting | Binary A only | Any A | Any A |
| Theory | Consistency | + Normality | + Efficiency |
| Assumptions | Strong | Medium | Weaker |
| Computation | Fast | Slow | Medium |
| Software | R package | None | R + Python |
This skill works with:
Version: 1.0 Created: 2025-12-08 Domain: Research Strategy, Literature Review