SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
原文の言語: 英語
メニュー
SkillsMP は yogsoth-ai/knowledge-acquisition から 77 件の skill を収集しています。skill を開くとソースと詳細を確認できます。
収集済み skill 77 件中 40 件を表示しています。
SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
原文の言語: 英語
Assess systematic biases in the evidence body — publication bias, reporting bias, and selective outcome reporting. Budget: 40 studies, 40 effect sizes, 40 web searches.
原文の言語: 英語
Track evidence accumulation over time — cumulative meta-analysis protocol design. Budget: 40 studies, 40 effect sizes, 30 web searches.
原文の言語: 英語
Design structured data extraction form for systematic meta-analysis data collection
原文の言語: 英語
Systematically extract effect sizes and conditions from papers for meta-analytic synthesis
原文の言語: 英語
Determine effect size types and calculation methods for meta-analytic synthesis
原文の言語: 英語
Build evidence network graph for network meta-analysis — nodes, edges, geometry assessment
原文の言語: 英語
Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting
原文の言語: 英語
Explain why different studies reach different conclusions — heterogeneity investigation protocol. Budget: 30 studies, 30 effect sizes, 50 web searches.
原文の言語: 英語
Identify and classify sources of between-study heterogeneity (clinical, methodological, statistical)
原文の言語: 英語
Define inclusion/exclusion criteria for systematic study selection in meta-analysis
原文の言語: 英語
Cross-Study Statistical Synthesis Campaign — 5 strategies for systematic collection and methodological planning of multi-study evidence synthesis. Covers pairwise, network, cumulative meta-analysis, heterogeneity investigation, and bias detection. Stops at…
原文の言語: 英語
Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol
原文の言語: 英語
Compare N methods simultaneously including indirect evidence — network meta-analysis protocol design. Budget: 50 studies, 80 effect sizes, 60 web searches.
原文の言語: 英語
Compare two methods across multiple studies — paired meta-analysis protocol design. Budget: 30 studies, 30 effect sizes, 40 web searches.
原文の言語: 英語
Construct PICO/PECO framework for the meta-analysis research question
原文の言語: 英語
Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias
原文の言語: 英語
Methodological quality and bias risk assessment of included studies using validated tools
原文の言語: 英語
Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools
原文の言語: 英語
Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks
原文の言語: 英語
Detect annotation artifacts and shortcuts in benchmarks
原文の言語: 英語
Evaluation Methodology Archaeology Campaign — 5 strategies for systematic analysis of AI/ML benchmarks, metrics, and leaderboards. Reveals construct validity issues, saturation, data contamination, and evaluation protocol inconsistencies.
原文の言語: 英語
Systematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches
原文の言語: 英語
Identify and catalog all relevant benchmarks in target domain
原文の言語: 英語
Produce final structured audit report
原文の言語: 英語
Build capability taxonomy, map existing benchmark coverage
原文の言語: 英語
Evaluate whether benchmark measures its claimed capability
原文の言語: 英語
Detect train-test data leakage and memorization artifacts
原文の言語: 英語
Map evaluation coverage, identify untested capability dimensions — 20 benchmarks, 30 papers, 50 web searches
原文の言語: 英語
Assess documentation completeness against BetterBench/Datasheets standards
原文の言語: 英語
Compare implementation differences of same benchmark across papers
原文の言語: 英語
Analyze leaderboard score distributions, compression, selective reporting
原文の言語: 英語
Decompose composite metrics into constituent signals, analyze polarity and ceiling effects
原文の言語: 英語
Extract evaluation protocol parameters from papers
原文の言語: 英語
Analyze evaluation protocol differences across papers for same benchmark — 5 benchmarks, 60 papers, 30 web searches
原文の言語: 英語
Track score trajectories, detect saturation/failure points — 15 benchmarks, 50 papers, 60 web searches
原文の言語: 英語
Collect historical scores, fit saturation curves, detect inflection points
原文の言語: 英語
Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches
原文の言語: 英語
Patent claim syntax parsing — independent/dependent relationships and element extraction
原文の言語: 英語
Determine patent legal status — active, expired, pending, lapsed, or revoked
原文の言語: 英語