SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
원문 언어: 영어
메뉴
SkillsMP는 yogsoth-ai/knowledge-acquisition에서 77개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.
수집된 skill 77개 중 40개를 표시합니다.
SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
원문 언어: 영어
Assess systematic biases in the evidence body — publication bias, reporting bias, and selective outcome reporting. Budget: 40 studies, 40 effect sizes, 40 web searches.
원문 언어: 영어
Track evidence accumulation over time — cumulative meta-analysis protocol design. Budget: 40 studies, 40 effect sizes, 30 web searches.
원문 언어: 영어
Design structured data extraction form for systematic meta-analysis data collection
원문 언어: 영어
Systematically extract effect sizes and conditions from papers for meta-analytic synthesis
원문 언어: 영어
Determine effect size types and calculation methods for meta-analytic synthesis
원문 언어: 영어
Build evidence network graph for network meta-analysis — nodes, edges, geometry assessment
원문 언어: 영어
Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting
원문 언어: 영어
Explain why different studies reach different conclusions — heterogeneity investigation protocol. Budget: 30 studies, 30 effect sizes, 50 web searches.
원문 언어: 영어
Identify and classify sources of between-study heterogeneity (clinical, methodological, statistical)
원문 언어: 영어
Define inclusion/exclusion criteria for systematic study selection in meta-analysis
원문 언어: 영어
Cross-Study Statistical Synthesis Campaign — 5 strategies for systematic collection and methodological planning of multi-study evidence synthesis. Covers pairwise, network, cumulative meta-analysis, heterogeneity investigation, and bias detection. Stops at…
원문 언어: 영어
Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol
원문 언어: 영어
Compare N methods simultaneously including indirect evidence — network meta-analysis protocol design. Budget: 50 studies, 80 effect sizes, 60 web searches.
원문 언어: 영어
Compare two methods across multiple studies — paired meta-analysis protocol design. Budget: 30 studies, 30 effect sizes, 40 web searches.
원문 언어: 영어
Construct PICO/PECO framework for the meta-analysis research question
원문 언어: 영어
Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias
원문 언어: 영어
Methodological quality and bias risk assessment of included studies using validated tools
원문 언어: 영어
Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools
원문 언어: 영어
Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks
원문 언어: 영어
Detect annotation artifacts and shortcuts in benchmarks
원문 언어: 영어
Evaluation Methodology Archaeology Campaign — 5 strategies for systematic analysis of AI/ML benchmarks, metrics, and leaderboards. Reveals construct validity issues, saturation, data contamination, and evaluation protocol inconsistencies.
원문 언어: 영어
Systematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches
원문 언어: 영어
Identify and catalog all relevant benchmarks in target domain
원문 언어: 영어
Produce final structured audit report
원문 언어: 영어
Build capability taxonomy, map existing benchmark coverage
원문 언어: 영어
Evaluate whether benchmark measures its claimed capability
원문 언어: 영어
Detect train-test data leakage and memorization artifacts
원문 언어: 영어
Map evaluation coverage, identify untested capability dimensions — 20 benchmarks, 30 papers, 50 web searches
원문 언어: 영어
Assess documentation completeness against BetterBench/Datasheets standards
원문 언어: 영어
Compare implementation differences of same benchmark across papers
원문 언어: 영어
Analyze leaderboard score distributions, compression, selective reporting
원문 언어: 영어
Decompose composite metrics into constituent signals, analyze polarity and ceiling effects
원문 언어: 영어
Extract evaluation protocol parameters from papers
원문 언어: 영어
Analyze evaluation protocol differences across papers for same benchmark — 5 benchmarks, 60 papers, 30 web searches
원문 언어: 영어
Track score trajectories, detect saturation/failure points — 15 benchmarks, 50 papers, 60 web searches
원문 언어: 영어
Collect historical scores, fit saturation curves, detect inflection points
원문 언어: 영어
Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches
원문 언어: 영어
Patent claim syntax parsing — independent/dependent relationships and element extraction
원문 언어: 영어
Determine patent legal status — active, expired, pending, lapsed, or revoked
원문 언어: 영어