Skip to main content

aris-autonomous-ml-research

Use ARIS (Auto-Research-In-Sleep) for autonomous ML research — idea generation, paper review, experiment automation, and cross-model collaboration with Claude Code, Codex, or any LLM agent.

Ir para a instalação

Informações da origem

Repositório
reason-machines/claude-code-skills
Última atividade na origem
16 de maio de 2026 às 16:22
Idioma detectado do SKILL.md
inglês
Estrelas
4
Forks
1

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
aris-autonomous-ml-research
description
Use ARIS (Auto-Research-In-Sleep) for autonomous ML research — idea generation, paper review, experiment automation, and cross-model collaboration with Claude Code, Codex, or any LLM agent.
triggers
["set up ARIS for autonomous research","use ARIS to generate research ideas","run autonomous ML experiments with ARIS","review a paper using ARIS cross-model workflow","configure ARIS research pipeline","help me with ARIS research automation","write a rebuttal using ARIS","start an ARIS research workflow"]
# ARIS Autonomous ML Research > Skill by [ara.so](https://ara.so) — Claude Code Skills collection. ARIS (Auto-Research-In-Sleep) is a lightweight Markdown-based system for autonomous ML research that orchestrates cross-model collaboration. It enables AI agents to discover ideas, review papers, run experiments, and write rebuttals — all autonomously. Works with Claude Code, Codex CLI, Cursor, Trae, Antigravity, or any LLM agent. ## What ARIS Does - **Idea Generation**: Automatically discovers research ideas from arXiv papers, GitHub repos, or research directions - **Cross-Model Review**: Uses different LLMs for execution vs. review to break self-play blind spots (e.g., Claude Code executes, GPT-5.4 reviews) - **Experiment Automation**: Clones codebases, runs experiments, analyzes results - **Paper Writing**: Generates drafts with auto-review loops to improve quality - **Rebuttal Generation**: Parses reviews, builds strategy, drafts rebuttals under character limits - **Research Wiki**: Persistent knowledge base tracking papers, ideas, experiments, and claims ## Installation ### As Claude Code Skills (In-Editor) ```bash # Clone skills into your Claude Code skills directory cd ~/claude-code-skills # or your skills path git clone https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git aris cd aris/skills # Skills are now available in Claude Code ``` ### As Standalone CLI (ARIS-Code) ```bash # Download latest release # Visit: https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep/releases/latest # Or install from source git clone https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git cd Auto-claude-code-research-in-sleep/aris-code cargo build --release # Run setup ./target/release/aris-code # Follow interactive setup to configure API keys ``` ### Environment Setup ARIS requires API keys for LLM providers: ```bash # For Claude (primary executor) export ANTHROPIC_API_KEY=your_key_here # For OpenAI (reviewer/alternative) export OPENAI_API_KEY=your_key_here # For alternative Chinese models (optional) export MOONSHOT_API_KEY=your_key_here # Kimi export MINIMAX_API_KEY=your_key_here # MiniMax export GLM_API_KEY=your_key_here # GLM export DEEPSEEK_API_KEY=your_key_here # DeepSeek export DOUBAO_API_KEY=your_key_here # Doubao ``` ## Key Commands ### Standalone CLI Commands ```bash # Interactive setup /setup # Run full research pipeline /research-pipeline "your research direction" # Run with reference paper and base repo /research-pipeline "improve method X" -- ref paper: https://arxiv.org/abs/2406.04329, base repo: https://github.com/org/project # Generate rebuttal from reviews /rebuttal "paper/ + reviews" -- venue: ICML, character limit: 5000 # Plan mode (step-by-step breakdown) /plan "implement transformer variant" # Research Wiki operations /wiki add paper <arxiv_url> /wiki add idea "your idea description" /wiki query "search term" /wiki export # Meta-optimization (self-improvement) /meta-optimize # Task management /tasks /tasks add "task description" /tasks complete <id> # Help and info /help /models # List available models ``` ### Claude Code Skill Commands When using ARIS as Claude Code skills, trigger workflows with natural language: ``` "Use ARIS to review this paper: https://arxiv.org/abs/2406.04329" "Generate research ideas about discrete diffusion models" "Run the research pipeline for improving attention mechanisms" "Help me write a rebuttal for these ICML reviews" ``` ## Core Workflows ### 1. Full Research Pipeline End-to-end autonomous research: ```bash /research-pipeline "factorized gap in discrete diffusion LMs" ``` **What happens:** 1. Discovers related papers from arXiv 2. Generates novel research ideas 3. Reviews ideas with external LLM (cross-model) 4. Runs experiments on top ideas 5. Writes paper draft 6. Auto-reviews and improves draft 7. Outputs final paper + code **With reference paper + codebase:** ```bash /research-pipeline "improve attention efficiency" -- ref paper: https://arxiv.org/abs/2305.xxxx, base repo: https://github.com/org/attention-impl ``` ARIS reads the paper → finds weaknesses → uses that specific codebase → generates targeted improvements. ### 2. Targeted Idea Discovery Generate ideas from specific sources: ```python # In Claude Code, reference the skill """ Use Workflow 1: DiscoverPaper skill to find ideas from: - arXiv search: "vision transformers" - GitHub repo: https://github.com/google-research/vision_transformer - Local paper: ./papers/vit_analysis.pdf """ ``` **ARIS will:** - Parse papers/code - Extract key insights - Generate 5-10 novel ideas - Route to external reviewer - Return scored, critiqued ideas ### 3. Paper Review Loop Multi-round review with automated improvements: ```bash # Standalone CLI /review-paper paper_draft.md --rounds 3 # In Claude Code "Review this draft and improve it through 3 rounds: ./draft.md" ``` **Review process:** 1. External LLM critiques (e.g., GPT-5.4) 2. Claude Code addresses weaknesses 3. Repeat until score plateau or max rounds 4. Final output with score progression graph ### 4. Rebuttal Generation Parse reviews and draft rebuttal: ```bash /rebuttal "paper_dir/" -- venue: ICML, character limit: 5000 ``` **Phases:** 1. **Parse reviews**: Extract all reviewer concerns 2. **Build strategy**: Map concerns → responses 3. **Draft rebuttal**: Generate structured response 4. **Format check**: Ensure under character limit **Quick mode** (stop before drafting): ```bash /rebuttal "paper_dir/" -- venue: NeurIPS, character limit: 8000, quick mode: true ``` ## Configuration ### Model Selection ARIS supports multiple executor + reviewer combinations: **In standalone CLI:** ```bash # Interactive setup /setup # Choose: # 1. Claude (Anthropic) # 2. OpenAI (GPT-4/5) # 3. Kimi (Moonshot) # 4. MiniMax # 5. GLM (Zhipu) # 6. DeepSeek # 7. Doubao # 8. LM Studio (local) # ... and more ``` **In skill files (YAML frontmatter):** ```yaml --- executor: claude-opus-4.7 # Primary LLM reviewer: gpt-5.5 # Review LLM --- ``` **Available executors:** - `claude-opus-4.7`, `claude-sonnet-4.5` - `gpt-5.5`, `gpt-5.4`, `o1`, `o3`, `o4` - `kimi-k2.5`, `kimi-k3` - `minimax-m2.7`, `minimax-pro` - `glm-5`, `glm-5-plus` - `deepseek-v3`, `deepseek-r1` - `doubao-lite`, `doubao-pro` **Reviewer routing:** ```yaml reviewer: oracle-pro # GPT-5.4 Pro via Oracle MCP (strongest) reviewer: gpt-5.4 # GPT-5.4 standard reviewer: claude-opus # Claude for review reviewer: auto # Smart routing based on executor ``` ### Research Wiki Configuration Enable persistent memory across sessions: ```bash # In CLI /setup # Enable "Research Wiki" option # Or set in config wiki: enabled: true path: ~/.aris/wiki/ auto_commit: true # Git commit after each change ``` **Wiki structure:** ``` ~/.aris/wiki/ ├── papers/ # Tracked papers ├── ideas/ # Research ideas ├── experiments/ # Experiment results ├── claims/ # Key claims and evidence └── graph.json # Relationship graph ``` ### Proxy and Custom Endpoints **HTTP/HTTPS proxy:** ```bash export HTTP_PROXY=http://localhost:8080 export HTTPS_PROXY=http://localhost:8080 # Or in /setup # Select provider → Configure proxy URL ``` **Custom API endpoints:** ```bash # For Anthropic-compatible proxies (Bedrock, etc.) /setup # Provider: Anthropic # Custom base URL: https://bedrock.amazonaws.com/anthropic/ ``` **Local models (LM Studio/Ollama):** ```bash /setup # Select "LM Studio / Ollama (Local)" # Base URL: http://localhost:1234/v1 # Model: local-model-name ``` ## Code Examples ### Example 1: Idea Discovery from Paper ```python # skill: workflow-1-discover-paper.md """ I want to discover ideas from this paper: https://arxiv.org/abs/2305.14342 (Transformer-XL) Focus on: attention mechanism improvements """ # ARIS will: # 1. Download and parse paper # 2. Extract key insights about XL attention # 3. Generate 5-10 novel ideas # 4. Send to GPT-5.4 for review # 5. Return scored ideas with critiques ``` **Output structure:** ```markdown ## Discovered Ideas (Reviewed) ### Idea 1: Factorized Relative Position Embeddings (Score: 8.5/10) **Core insight:** XL uses dense relative position matrix — factorize it. **Reviewer critique (GPT-5.4):** - ✅ Novelty: High (not explored in XL paper) - ✅ Feasibility: Doable (standard tensor decomposition) - ⚠️ Impact: Need to verify on long sequences - ⚠️ Risk: May hurt performance if rank too low **Next steps:** Implement SVD-based factorization, benchmark on PG-19 --- ### Idea 2: Learnable Decay for Relative Attention (Score: 7.2/10) ... ``` ### Example 2: Experiment Automation ```python # skill: workflow-2-run-experiment.md """ Clone https://github.com/kimiyoung/transformer-xl Implement Idea 1 (factorized position embeddings) Run on enwik8 benchmark Compare with baseline """ # ARIS will: # 1. Clone repo # 2. Create experiment branch # 3. Modify model code (e.g., pytorch_modules/rel_multihead_attn.py) # 4. Set up training config # 5. Run experiment # 6. Parse results # 7. Generate comparison report ``` **Generated experiment code:** ```python # aris_experiments/factorized_rel_pos/model_patch.py import torch import torch.nn as nn class FactorizedRelativeAttention(nn.Module): def __init__(self, d_model, n_heads, rank=64): super().__init__() self.d_model = d_model self.n_heads = n_heads self.rank = rank # Factorized position embeddings: (seq_len, d_model) ≈ U @ V^T self.U = nn.Parameter(torch.randn(2048, rank)) # max_seq_len self.V = nn.Parameter(torch.randn(d_model, rank)) def forward(self, q, k, v, pos_emb): # Compute relative position on-the-fly rel_pos = self.U @ self.V.t() # (2048, d_model) # ... rest of attention logic ``` **Experiment results:** ```markdown ## Experiment Results: Factorized Relative Position | Metric | Baseline (XL) | Ours (Factorized) | Δ | |--------|---------------|-------------------|---| | PPL (enwik8) | 1.06 | 1.08 | +0.02 ↓ | | Speed (tok/s) | 12.3k | 18.7k | +52% ↑ | | Memory (GB) | 11.2 | 7.8 | -30% ↑ | | Params (M) | 277 | 261 | -5.8% ↑ | ✅ **Success**: 52% faster, 30% less memory, minor PPL degradation acceptable for long-context tasks. ``` ### Example 3: Paper Writing with Auto-Review ```python # skill: workflow-3-write-paper.md """ Write a paper about the factorized position embeddings experiment. Title: "Efficient Transformers via Factorized Relative Attention" Target venue: ICML 2026 Sections: Abstract, Introduction, Method, Experiments, Conclusion """ # ARIS will: # 1. Generate initial draft # 2. Send to GPT-5.4 for review # 3. Address critiques (e.g., "add ablation study") # 4. Re-review # 5. Repeat for N rounds or until score plateau ``` **Review loop:** ```markdown ## Round 1 Review (GPT-5.4, Score: 6.5/10) **Strengths:** - Clear motivation (long-context efficiency) - Solid experimental results **Weaknesses:** - Missing ablation on rank hyperparameter - No comparison with Linear Attention baseline - Introduction lacks related work on efficient Transformers **Suggestions:** 1. Add Table 2: Rank ablation (r=16,32,64,128) 2. Cite Performer, Linformer in related work 3. Add wall-clock time comparison --- ## Round 2 Review (GPT-5.4, Score: 8.1/10) **Improvements:** ✅ Added rank ablation (Table 2) ✅ Expanded related work ✅ Wall-clock benchmarks included **Remaining issues:** - Figure 3 caption unclear - Conclusion should mention future work --- ## Round 3 Review (GPT-5.4, Score: 8.8/10) **Near-ready:** Minor edits only. Ready for submission. ``` ### Example 4: Rebuttal Generation ```bash
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub