Skip to main content

claude-code-local-mlx

Run Claude Code 100% on-device with local AI on Apple Silicon using MLX-native models (Qwen 3.5 122B, Llama 3.3 70B, Gemma 4 31B, DeepSeek V4 Flash)

설치로 이동

소스 정보

저장소
reason-machines/claude-code-skills
최근 소스 활동
2026년 5월 17일 03:08
감지된 SKILL.md 언어
영어
스타
4
포크
1

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
claude-code-local-mlx
description
Run Claude Code 100% on-device with local AI on Apple Silicon using MLX-native models (Qwen 3.5 122B, Llama 3.3 70B, Gemma 4 31B, DeepSeek V4 Flash)
triggers
["set up local claude code","run claude code offline","install mlx local ai","configure qwen llama gemma locally","run deepseek v4 flash locally","setup airgap ai on apple silicon","run anthropic api server locally","configure mlx-lm with claude code"]
# Claude Code Local — MLX On-Device AI > Skill by [ara.so](https://ara.so) — Claude Code Skills collection. Run Claude Code 100% on-device with local AI on Apple Silicon. This project provides an MLX-native Anthropic-API-compatible server that lets you use powerful local models (Qwen 3.5 122B at 65 tok/s, Llama 3.3 70B, Gemma 4 31B, DeepSeek V4 Flash with 1M context) as drop-in replacements for Claude. Built for privacy-critical workflows (NDA, legal, healthcare) where data cannot leave the device. ## What It Does - **MLX-native inference** optimized for Apple Silicon (M1/M2/M3/M4) - **Anthropic API compatibility** — point Claude Code clients at `localhost:8000` - **Four model options**: Gemma 4 31B (fast), Qwen 3.5 122B (balanced), Llama 3.3 70B (dense), DeepSeek V4 Flash (1M context) - **100% offline** — works in airplane mode, airgap-ready - **Voice mode** — hands-free coding with on-device STT/TTS - **Browser agent** — remote access from any device on your LAN ## Installation ### Prerequisites ```bash # macOS with Apple Silicon (M1/M2/M3/M4) # Minimum RAM: # - Gemma 4 31B: 32 GB # - Qwen 3.5 122B: 96 GB # - Llama 3.3 70B: 96 GB # - DeepSeek V4 Flash: 128 GB # Python 3.10+ python3 --version # Homebrew (for dependencies) /bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)" ``` ### Quick Start (3 Commands) ```bash # 1. Clone the repo git clone https://github.com/nicedreamzapp/claude-code-local.git cd claude-code-local # 2. Install dependencies pip install -r requirements.txt # 3. Start the server (Qwen 3.5 122B example) bash scripts/start-mlx-server.sh ``` The server will: 1. Download the model from HuggingFace (~65 GB for Qwen) 2. Start an Anthropic-API-compatible server on `http://localhost:8000` 3. Print usage instructions ### Model Selection Set `MLX_MODEL` environment variable before starting: ```bash # Gemma 4 31B (fast, 32GB+ RAM) export MLX_MODEL="bartowski/gemma-4-code-it-31b-MLX-4bit" bash scripts/start-mlx-server.sh # Qwen 3.5 122B (balanced, 96GB+ RAM) export MLX_MODEL="mlx-community/Qwen3.5-MoE-A10B-4bit" bash scripts/start-mlx-server.sh # Llama 3.3 70B (dense, 96GB+ RAM) export MLX_MODEL="divinetribe/Llama-3.3-70B-Instruct-abliterated-8bit-mlx" bash scripts/start-mlx-server.sh # DeepSeek V4 Flash (1M context, 128GB+ RAM) — uses ds4 engine # See DeepSeek section below ``` ### DeepSeek V4 Flash Setup DeepSeek uses the separate `ds4` engine: ```bash # 1. Install ds4 git clone https://github.com/antirez/ds4.git ~/.ds4 cd ~/.ds4 make # 2. Download model (81GB for q2, 153GB for q4) huggingface-cli download antirez/deepseek-v4-gguf \ deepseek-v4-q2.gguf --local-dir ~/.ds4/models # 3. Start ds4 server wrapper ~/.local/bin/ds4-server-up # 4. Use with Claude Code ~/.local/bin/claude-ds4 "your prompt here" ``` ## Configuration ### Claude Code Client Setup Point your Claude Code client to the local server: ```bash # Set API base URL export ANTHROPIC_API_KEY="sk-ant-local-placeholder" # Any string works export ANTHROPIC_BASE_URL="http://localhost:8000" # Or use the wrapper script ./claude-local "create a react component for a todo list" ``` ### Server Configuration Create `config/server.yaml`: ```yaml # MLX server config server: host: "0.0.0.0" # Listen on all interfaces port: 8000 model: name: "mlx-community/Qwen3.5-MoE-A10B-4bit" max_tokens: 256000 # Qwen supports 256K context temperature: 0.7 performance: cache_prompt: true # Enable prompt caching max_kv_size: 256000 ``` ### Launcher Scripts The project includes `.command` launchers for double-click execution: ```bash # macOS launchers (in repo root) "Claude Local.command" # Qwen 3.5 122B "Gemma 4 Code.command" # Gemma 4 31B "DeepSeek V4 Flash.app" # DeepSeek bundle ``` ## Key Commands ### Starting the Server ```bash # Basic start bash scripts/start-mlx-server.sh # With custom model MLX_MODEL="mlx-community/Qwen3.5-MoE-A10B-4bit" \ bash scripts/start-mlx-server.sh # With custom port PORT=8080 bash scripts/start-mlx-server.sh # Monitor resource usage python3 scripts/monitor.py ``` ### Using the API ```python import anthropic # Point to local server client = anthropic.Anthropic( api_key="sk-ant-local-placeholder", base_url="http://localhost:8000" ) # Stream a response with client.messages.stream( model="qwen-3.5-122b", # Name doesn't matter, uses running model max_tokens=4096, messages=[ {"role": "user", "content": "Write a Python function to parse JSON"} ] ) as stream: for text in stream.text_stream: print(text, end="", flush=True) ``` ### Voice Mode Setup ```bash # Install voice dependencies pip install -r requirements-voice.txt # Download voice models (whisper + kokoro TTS) python3 scripts/download-voice-models.py # Start voice-enabled server bash scripts/start-voice-server.sh # Talk to Claude Code (hands-free) python3 scripts/voice-client.py ``` ### Browser Agent Access Claude Code from any device on your LAN: ```bash # Start browser agent server python3 scripts/browser-agent-server.py # Open in any browser (phone, tablet, etc.) # Navigate to: http://YOUR_MAC_IP:5000 ``` ## Real Code Examples ### Example 1: Basic Coding Task ```python #!/usr/bin/env python3 """Send a coding task to local Claude Code""" import anthropic client = anthropic.Anthropic( api_key="sk-ant-local", base_url="http://localhost:8000" ) response = client.messages.create( model="local", max_tokens=8192, messages=[{ "role": "user", "content": """Create a FastAPI endpoint that: 1. Accepts POST with JSON {query: str} 2. Validates query is 1-500 chars 3. Returns {result: str, timestamp: int} Include error handling and type hints.""" }] ) print(response.content[0].text) ``` ### Example 2: Document Analysis (Airgap Mode) ```python #!/usr/bin/env python3 """Analyze confidential document with offline AI""" import anthropic import sys def analyze_nda(filepath): with open(filepath, 'r') as f: doc_text = f.read() client = anthropic.Anthropic( api_key="sk-ant-local", base_url="http://localhost:8000" ) response = client.messages.create( model="llama-3.3-70b", max_tokens=16384, messages=[{ "role": "user", "content": f"""Analyze this NDA for: 1. Term length 2. Geographic scope 3. Definition of confidential info 4. Carve-outs and exceptions 5. Red flags Document: {doc_text}""" }] ) return response.content[0].text if __name__ == "__main__": result = analyze_nda(sys.argv[1]) print(result) ``` ### Example 3: Long-Context Processing (DeepSeek) ```python #!/usr/bin/env python3 """Use DeepSeek V4 Flash for 1M token context""" import anthropic import glob def analyze_codebase(directory): # Collect all Python files files = glob.glob(f"{directory}/**/*.py", recursive=True) codebase = "" for filepath in files[:100]: # Process up to 100 files with open(filepath, 'r') as f: codebase += f"\n\n# {filepath}\n{f.read()}" client = anthropic.Anthropic( api_key="sk-ant-local", base_url="http://localhost:8001" # ds4 runs on 8001 ) response = client.messages.create( model="deepseek-v4-flash", max_tokens=32000, messages=[{ "role": "user", "content": f"""Analyze this entire codebase: 1. Architecture patterns used 2. Security vulnerabilities 3. Performance bottlenecks 4. Suggested refactorings Codebase: {codebase}""" }] ) return response.content[0].text result = analyze_codebase("./my-project") print(result) ``` ### Example 4: Streaming with Tool Use ```python #!/usr/bin/env python3 """Stream responses with tool calling""" import anthropic tools = [{ "name": "run_bash", "description": "Execute a bash command and return output", "input_schema": { "type": "object", "properties": { "command": {"type": "string", "description": "Bash command"} }, "required": ["command"] } }] client = anthropic.Anthropic( api_key="sk-ant-local", base_url="http://localhost:8000" ) with client.messages.stream( model="qwen-3.5", max_tokens=4096, tools=tools, messages=[{ "role": "user", "content": "List all Python files in current directory and count lines" }] ) as stream: for event in stream: if event.type == "content_block_delta": if hasattr(event.delta, "text"): print(event.delta.text, end="", flush=True) elif event.type == "tool_use": print(f"\n[Tool call: {event.name}]") ``` ## Common Patterns ### Pattern 1: Multi-Model Workflow ```python #!/usr/bin/env python3 """Use different models for different tasks""" import anthropic def get_client(model_port): return anthropic.Anthropic( api_key="sk-ant-local", base_url=f"http://localhost:{model_port}" ) # Fast model for quick responses (Gemma on 8000) gemma = get_client(8000) quick_response = gemma.messages.create( model="gemma-4", max_tokens=1024, messages=[{"role": "user", "content": "Summarize: <long text>"}] ) # Heavy model for complex reasoning (Qwen on 8001) qwen = get_client(8001) deep_analysis = qwen.messages.create( model="qwen-3.5", max_tokens=8192, messages=[{ "role": "user", "content": f"Analyze this summary: {quick_response.content[0].text}" }] ) ``` ### Pattern 2: Prompt Caching for Speed ```python #!/usr/bin/env python3 """Leverage KV cache for repeated contexts""" import anthropic client = anthropic.Anthropic( api_key="sk-ant-local", base_url="http://localhost:8000" ) # Long system prompt (cached after first use) system_prompt = """You are an expert Python developer. Follow PEP 8, use type hints, write docstrings. Prefer dataclasses over dicts, use pathlib over os.path. [... 5000 more tokens of coding guidelines ...]""" # First call: slow (prefills cache) response1 = client.messages.create( model="qwen", max_tokens=2048, system=system_prompt, messages=[{"role": "user", "content": "Write a parser"}] ) # Second call: fast (reuses cached prompt) response2 = client.messages.create( model="qwen", max_tokens=2048, system=system_prompt, # Same prompt = cache hit messages=[{"role": "user", "content": "Write a validator"}] ) ``` ### Pattern 3: Airgap Verification ```python #!/usr/bin/env python3 """Verify no network calls during inference""" import subprocess import anthropic import time # Start monitoring network monitor = subprocess.Popen( ["lsof", "-i", "-n", "-P"], stdout=subprocess.PIPE, stderr=subprocess.PIPE ) client = anthropic.Anthropic( api_key="sk-ant-local", base_url="http://localhost:8000" ) # Make AI request response = client.messages.create(
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기