Skip to main content

guidance

Constrain LLM output with grammars; guarantee valid JSON.

Quellinformationen

Repository
NousResearch/hermes-agent
Letzte Quellaktivität
24. Juli 2026 um 15:18
Erkannte Sprache von SKILL.md
Englisch
Sterne
249.930
Forks
53.284

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
4 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
guidance
description
Constrain LLM output with grammars; guarantee valid JSON.
version
1.0.1
author
Orchestra Research
license
MIT
dependencies
["guidance","transformers"]
platforms
["linux","macos","windows"]
metadata
{"hermes":{"tags":["Prompt Engineering","Guidance","Constrained Generation","Structured Output","JSON Validation","Grammar","Microsoft Research","Format Enforcement","Multi-Step Workflows"]}}
# Guidance: Constrained LLM Generation ## When to Use This Skill Use Guidance when you need to: - **Control LLM output syntax** with regex or grammars - **Guarantee valid JSON/XML/code** generation - **Reduce latency** vs traditional prompting approaches - **Enforce structured formats** (dates, emails, IDs, etc.) - **Build multi-step workflows** with Pythonic control flow - **Prevent invalid outputs** through grammatical constraints **GitHub Stars**: 18,000+ | **From**: Microsoft Research ## Installation ```bash # Base installation pip install guidance # With specific backends pip install guidance[transformers] # Hugging Face models pip install guidance[llama_cpp] # llama.cpp models ``` ## Quick Start ### Basic Example: Structured Generation ```python from guidance import models, gen # Load model (supports OpenAI, Transformers, llama.cpp) lm = models.OpenAI("gpt-4") # Generate with constraints result = lm + "The capital of France is " + gen("capital", max_tokens=5) print(result["capital"]) # "Paris" ``` ### Chat format with a local model > **Constraint support requires local logit access.** Regex, `select()`, and > grammar-based constrained generation only work with local backends > (`Transformers`, `LlamaCpp`). Remote API backends (`OpenAI`, and Azure > variants) support unconstrained `gen()` / chat only — they cannot enforce > token-level constraints. guidance 0.3.x has no `models.Anthropic` class. ```python from guidance import models, gen, system, user, assistant # Local model (supports constrained generation) lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Use context managers for chat format with system(): lm += "You are a helpful assistant." with user(): lm += "What is the capital of France?" with assistant(): lm += gen(max_tokens=20) ``` ## Core Concepts ### 1. Context Managers Guidance uses Pythonic context managers for chat-style interactions. ```python from guidance import system, user, assistant, gen lm = models.Transformers("microsoft/Phi-4-mini-instruct") # System message with system(): lm += "You are a JSON generation expert." # User message with user(): lm += "Generate a person object with name and age." # Assistant response with assistant(): lm += gen("response", max_tokens=100) print(lm["response"]) ``` **Benefits:** - Natural chat flow - Clear role separation - Easy to read and maintain ### 2. Constrained Generation Guidance ensures outputs match specified patterns using regex or grammars. #### Regex Constraints ```python from guidance import models, gen lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Constrain to valid email format lm += "Email: " + gen("email", regex=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}") # Constrain to date format (YYYY-MM-DD) lm += "Date: " + gen("date", regex=r"\d{4}-\d{2}-\d{2}") # Constrain to phone number lm += "Phone: " + gen("phone", regex=r"\d{3}-\d{3}-\d{4}") print(lm["email"]) # Guaranteed valid email print(lm["date"]) # Guaranteed YYYY-MM-DD format ``` **How it works:** - Regex converted to grammar at token level - Invalid tokens filtered during generation - Model can only produce matching outputs #### Selection Constraints ```python from guidance import models, gen, select lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Constrain to specific choices lm += "Sentiment: " + select(["positive", "negative", "neutral"], name="sentiment") # Multiple-choice selection lm += "Best answer: " + select( ["A) Paris", "B) London", "C) Berlin", "D) Madrid"], name="answer" ) print(lm["sentiment"]) # One of: positive, negative, neutral print(lm["answer"]) # One of: A, B, C, or D ``` ### 3. Token Healing Guidance automatically "heals" token boundaries between prompt and generation. **Problem:** Tokenization creates unnatural boundaries. ```python # Without token healing prompt = "The capital of France is " # Last token: " is " # First generated token might be " Par" (with leading space) # Result: "The capital of France is Paris" (double space!) ``` **Solution:** Guidance backs up one token and regenerates. ```python from guidance import models, gen lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Token healing enabled by default lm += "The capital of France is " + gen("capital", max_tokens=5) # Result: "The capital of France is Paris" (correct spacing) ``` **Benefits:** - Natural text boundaries - No awkward spacing issues - Better model performance (sees natural token sequences) ### 4. Grammar-Based Generation Define complex structures by composing grammar functions. The template-string `grammar=` form is not part of current guidance — build grammars from composable functions, or use `guidance.json()` for JSON. ```python from guidance import models, gen from guidance import json as gen_json from pydantic import BaseModel, Field lm = models.Transformers("microsoft/Phi-4-mini-instruct") # JSON via a Pydantic schema (guidance.json compiles the schema to a grammar) class Person(BaseModel): name: str = Field(pattern=r"[A-Za-z ]+") age: int email: str = Field(pattern=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}") lm += gen_json(name="person", schema=Person) print(lm["person"]) # Guaranteed valid JSON matching the schema # Or compose grammar functions directly: grammar = "name=" + gen("name", regex=r"[A-Za-z ]+") + " age=" + gen("age", regex=r"[0-9]+") lm += grammar ``` **Use cases:** - Complex structured outputs - Nested data structures - Programming language syntax - Domain-specific languages ### 5. Guidance Functions Create reusable generation patterns with the `@guidance` decorator. ```python from guidance import guidance, gen, models @guidance def generate_person(lm): """Generate a person with name and age.""" lm += "Name: " + gen("name", max_tokens=20, stop="\n") lm += "\nAge: " + gen("age", regex=r"[0-9]+", max_tokens=3) return lm # Use the function lm = models.Transformers("microsoft/Phi-4-mini-instruct") lm = generate_person(lm) print(lm["name"]) print(lm["age"]) ``` **Stateful Functions:** ```python @guidance(stateless=False) def react_agent(lm, question, tools, max_rounds=5): """ReAct agent with tool use.""" lm += f"Question: {question}\n\n" for i in range(max_rounds): # Thought lm += f"Thought {i+1}: " + gen("thought", stop="\n") # Action lm += "\nAction: " + select(list(tools.keys()), name="action") # Execute tool tool_result = tools[lm["action"]]() lm += f"\nObservation: {tool_result}\n\n" # Check if done lm += "Done? " + select(["Yes", "No"], name="done") if lm["done"] == "Yes": break # Final answer lm += "\nFinal Answer: " + gen("answer", max_tokens=100) return lm ``` ## Backend Configuration ### OpenAI (remote — unconstrained only) > Remote API backends cannot do constrained generation (regex/select/grammar); > use them only for plain chat/`gen()`. For constraints, use a local backend. ```python from guidance import models lm = models.OpenAI( model="gpt-4o-mini", api_key="your-api-key" # Or set OPENAI_API_KEY env var ) ``` ### Local Models (Transformers) ```python from guidance.models import Transformers lm = Transformers( "microsoft/Phi-4-mini-instruct", device="cuda" # Or "cpu" ) ``` ### Local Models (llama.cpp) ```python from guidance.models import LlamaCpp lm = LlamaCpp( model_path="/path/to/model.gguf", n_ctx=4096, n_gpu_layers=35 ) ``` ## Common Patterns ### Pattern 1: JSON Generation ```python from guidance import models, gen, system, user, assistant lm = models.Transformers("microsoft/Phi-4-mini-instruct") with system(): lm += "You generate valid JSON." with user(): lm += "Generate a user profile with name, age, and email." with assistant(): lm += """{ "name": """ + gen("name", regex=r'"[A-Za-z ]+"', max_tokens=30) + """, "age": """ + gen("age", regex=r"[0-9]+", max_tokens=3) + """, "email": """ + gen("email", regex=r'"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}"', max_tokens=50) + """ }""" print(lm) # Valid JSON guaranteed ``` ### Pattern 2: Classification ```python from guidance import models, gen, select lm = models.Transformers("microsoft/Phi-4-mini-instruct") text = "This product is amazing! I love it." lm += f"Text: {text}\n" lm += "Sentiment: " + select(["positive", "negative", "neutral"], name="sentiment") lm += "\nConfidence: " + gen("confidence", regex=r"[0-9]+", max_tokens=3) + "%" print(f"Sentiment: {lm['sentiment']}") print(f"Confidence: {lm['confidence']}%") ``` ### Pattern 3: Multi-Step Reasoning ```python from guidance import models, gen, guidance @guidance def chain_of_thought(lm, question): """Generate answer with step-by-step reasoning.""" lm += f"Question: {question}\n\n" # Generate multiple reasoning steps for i in range(3): lm += f"Step {i+1}: " + gen(f"step_{i+1}", stop="\n", max_tokens=100) + "\n" # Final answer lm += "\nTherefore, the answer is: " + gen("answer", max_tokens=50) return lm lm = models.Transformers("microsoft/Phi-4-mini-instruct") lm = chain_of_thought(lm, "What is 15% of 200?") print(lm["answer"]) ``` ### Pattern 4: ReAct Agent ```python from guidance import models, gen, select, guidance @guidance(stateless=False) def react_agent(lm, question): """ReAct agent with tool use.""" tools = { "calculator": lambda expr: eval(expr), "search": lambda query: f"Search results for: {query}", } lm += f"Question: {question}\n\n" for round in range(5): # Thought lm += f"Thought: " + gen("thought", stop="\n") + "\n" # Action selection
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen