SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
Quellsprache: Englisch
Menü
SkillsMP hat 77 Skills aus yogsoth-ai/knowledge-acquisition gesammelt. Öffne einen Skill, um Quelle und Details zu prüfen.
Es werden 40 von 77 gesammelten Skills angezeigt.
SOTA Performance Baseline Campaign — 5 strategies for systematically collecting, standardizing, and analyzing performance data across methods. Produces standardized comparison tables, progress curves, and headroom analysis.
Quellsprache: Englisch
Assess systematic biases in the evidence body — publication bias, reporting bias, and selective outcome reporting. Budget: 40 studies, 40 effect sizes, 40 web searches.
Quellsprache: Englisch
Track evidence accumulation over time — cumulative meta-analysis protocol design. Budget: 40 studies, 40 effect sizes, 30 web searches.
Quellsprache: Englisch
Design structured data extraction form for systematic meta-analysis data collection
Quellsprache: Englisch
Systematically extract effect sizes and conditions from papers for meta-analytic synthesis
Quellsprache: Englisch
Determine effect size types and calculation methods for meta-analytic synthesis
Quellsprache: Englisch
Build evidence network graph for network meta-analysis — nodes, edges, geometry assessment
Quellsprache: Englisch
Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting
Quellsprache: Englisch
Explain why different studies reach different conclusions — heterogeneity investigation protocol. Budget: 30 studies, 30 effect sizes, 50 web searches.
Quellsprache: Englisch
Identify and classify sources of between-study heterogeneity (clinical, methodological, statistical)
Quellsprache: Englisch
Define inclusion/exclusion criteria for systematic study selection in meta-analysis
Quellsprache: Englisch
Cross-Study Statistical Synthesis Campaign — 5 strategies for systematic collection and methodological planning of multi-study evidence synthesis. Covers pairwise, network, cumulative meta-analysis, heterogeneity investigation, and bias detection. Stops at…
Quellsprache: Englisch
Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol
Quellsprache: Englisch
Compare N methods simultaneously including indirect evidence — network meta-analysis protocol design. Budget: 50 studies, 80 effect sizes, 60 web searches.
Quellsprache: Englisch
Compare two methods across multiple studies — paired meta-analysis protocol design. Budget: 30 studies, 30 effect sizes, 40 web searches.
Quellsprache: Englisch
Construct PICO/PECO framework for the meta-analysis research question
Quellsprache: Englisch
Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias
Quellsprache: Englisch
Methodological quality and bias risk assessment of included studies using validated tools
Quellsprache: Englisch
Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools
Quellsprache: Englisch
Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks
Quellsprache: Englisch
Detect annotation artifacts and shortcuts in benchmarks
Quellsprache: Englisch
Evaluation Methodology Archaeology Campaign — 5 strategies for systematic analysis of AI/ML benchmarks, metrics, and leaderboards. Reveals construct validity issues, saturation, data contamination, and evaluation protocol inconsistencies.
Quellsprache: Englisch
Systematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches
Quellsprache: Englisch
Identify and catalog all relevant benchmarks in target domain
Quellsprache: Englisch
Produce final structured audit report
Quellsprache: Englisch
Build capability taxonomy, map existing benchmark coverage
Quellsprache: Englisch
Evaluate whether benchmark measures its claimed capability
Quellsprache: Englisch
Detect train-test data leakage and memorization artifacts
Quellsprache: Englisch
Map evaluation coverage, identify untested capability dimensions — 20 benchmarks, 30 papers, 50 web searches
Quellsprache: Englisch
Assess documentation completeness against BetterBench/Datasheets standards
Quellsprache: Englisch
Compare implementation differences of same benchmark across papers
Quellsprache: Englisch
Analyze leaderboard score distributions, compression, selective reporting
Quellsprache: Englisch
Decompose composite metrics into constituent signals, analyze polarity and ceiling effects
Quellsprache: Englisch
Extract evaluation protocol parameters from papers
Quellsprache: Englisch
Analyze evaluation protocol differences across papers for same benchmark — 5 benchmarks, 60 papers, 30 web searches
Quellsprache: Englisch
Track score trajectories, detect saturation/failure points — 15 benchmarks, 50 papers, 60 web searches
Quellsprache: Englisch
Collect historical scores, fit saturation curves, detect inflection points
Quellsprache: Englisch
Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches
Quellsprache: Englisch
Patent claim syntax parsing — independent/dependent relationships and element extraction
Quellsprache: Englisch
Determine patent legal status — active, expired, pending, lapsed, or revoked
Quellsprache: Englisch