Structure ablations, statistical comparisons, and evidence tracing for Tenacious-Bench results. Use when an agent is preparing Delta A, Delta B, or Delta C comparisons, held-out scoring traces, confidence intervals, cost or latency comparisons, evidence…
Sanoy24/tenacious-bench
SkillsMP has collected 8 skills from Sanoy24/tenacious-bench. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 8
- GitHub stars
- 1
- GitHub forks
- 0
Skills in this repository
Showing 8 of 8 collected skills.
Build or extend Tenacious-Bench and its supporting Week 11 artifacts for the Sales Agent Evaluation Bench challenge. Use when an agent needs to turn Week 10 sales-agent traces, probes, style guidance, or public-signal inputs into a machine-verifiable…
Protect Tenacious-Bench split integrity and contamination resistance. Use when an agent is partitioning tasks, sealing held-out data, checking overlap between train and held_out, validating public-signal time windows, documenting contamination controls, or…
Create, expand, or review Tenacious-Bench tasks across the required authoring modes. Use when an agent is turning Week 10 traces, probe seeds, public-signal inputs, or sales artifacts into benchmark tasks, partition-ready records, metadata-rich JSON or JSONL…
Package Tenacious-Bench outputs into public, reviewer-friendly artifacts. Use when an agent is preparing a Hugging Face dataset or model card, datasheet, README, technical blog draft, publication checklist, executive memo support files, or…
Write, review, or refactor Python code for the Tenacious Bench project using clean, modern, and maintainable patterns. Use when an agent is adding or editing Python modules, CLI scripts, evaluators, dataset builders, contamination checks, training-data…
Design, tighten, or review machine-verifiable scoring rubrics for Tenacious-Bench tasks. Use when an agent is creating or revising schema fields, scoring evaluator logic, judge prompts, rubric dimensions, banned-phrase checks, grounding requirements, or…
Convert Tenacious-Bench artifacts into high-quality training inputs for Path A, Path B, or Path C. Use when an agent is formatting chat pairs, preference pairs, step-level labels, quality filters, or path-specific training partitions for LoRA, judge training,…