بنقرة واحدة
hermes-research-agent
يحتوي hermes-research-agent على 7 من skills المجمعة من ar0cket1، مع تغطية مهنية على مستوى المستودع وصفحات skill داخل الموقع.
Skills في هذا المستودع
Run LLM post-training on Tinker with CPU-side orchestration and remote GPU execution. Use when preparing or launching SFT, DPO, or PPO-style runs, checkpointing, sampling checkpoints, or resuming long-running jobs through Hermes Research Agent.
Run durable end-to-end LLM post-training research loops from zero-spec ideas to finished reports. Use when the goal is autonomous literature review, hypothesis selection, experiment planning, Tinker training, evaluation, checkpointing, and resumable long-running research work.
Plan and interpret model evaluations and ablations for post-training research. Use when comparing checkpoints, designing ablations, selecting benchmarks, or summarizing what changed after training.
Convert literature findings into concrete LLM experiment plans. Use when the agent has papers, blog posts, or benchmark notes and needs to turn them into hypotheses, training plans, dataset choices, and evaluation criteria.
Generate and score candidate LLM research ideas from sparse user goals. Use when a project starts with no concrete benchmark, dataset, or training recipe and the agent needs to propose plausible, testable directions.
Produce clear iteration memos and final research summaries from project artifacts. Use when converting experiment state, training outcomes, evaluation results, and open questions into reports for later review.
Expert guidance for fine-tuning LLMs with Axolotl - YAML configs, 100+ models, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, multimodal support