| name | guidance |
| description | Constrain LLM output with grammars; guarantee valid JSON. |
| version | 1.0.1 |
| author | Orchestra Research |
| license | MIT |
| dependencies | ["guidance","transformers"] |
| platforms | ["linux","macos","windows"] |
| metadata | {"hermes":{"tags":["Prompt Engineering","Guidance","Constrained Generation","Structured Output","JSON Validation","Grammar","Microsoft Research","Format Enforcement","Multi-Step Workflows"]}} |
Guidance: Constrained LLM Generation
When to Use This Skill
Use Guidance when you need to:
- Control LLM output syntax with regex or grammars
- Guarantee valid JSON/XML/code generation
- Reduce latency vs traditional prompting approaches
- Enforce structured formats (dates, emails, IDs, etc.)
- Build multi-step workflows with Pythonic control flow
- Prevent invalid outputs through grammatical constraints
GitHub Stars: 18,000+ | From: Microsoft Research
Installation
pip install guidance
pip install guidance[transformers]
pip install guidance[llama_cpp]
Quick Start
Basic Example: Structured Generation
from guidance import models, gen
lm = models.OpenAI("gpt-4")
result = lm + "The capital of France is " + gen("capital", max_tokens=5)
print(result["capital"])
Chat format with a local model
Constraint support requires local logit access. Regex, select(), and
grammar-based constrained generation only work with local backends
(Transformers, LlamaCpp). Remote API backends (OpenAI, and Azure
variants) support unconstrained gen() / chat only — they cannot enforce
token-level constraints. guidance 0.3.x has no models.Anthropic class.
from guidance import models, gen, system, user, assistant
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
with system():
lm += "You are a helpful assistant."
with user():
lm += "What is the capital of France?"
with assistant():
lm += gen(max_tokens=20)
Core Concepts
1. Context Managers
Guidance uses Pythonic context managers for chat-style interactions.
from guidance import system, user, assistant, gen
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
with system():
lm += "You are a JSON generation expert."
with user():
lm += "Generate a person object with name and age."
with assistant():
lm += gen("response", max_tokens=100)
print(lm["response"])
Benefits:
- Natural chat flow
- Clear role separation
- Easy to read and maintain
2. Constrained Generation
Guidance ensures outputs match specified patterns using regex or grammars.
Regex Constraints
from guidance import models, gen
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
lm += "Email: " + gen("email", regex=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}")
lm += "Date: " + gen("date", regex=r"\d{4}-\d{2}-\d{2}")
lm += "Phone: " + gen("phone", regex=r"\d{3}-\d{3}-\d{4}")
print(lm["email"])
print(lm["date"])
How it works:
- Regex converted to grammar at token level
- Invalid tokens filtered during generation
- Model can only produce matching outputs
Selection Constraints
from guidance import models, gen, select
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
lm += "Sentiment: " + select(["positive", "negative", "neutral"], name="sentiment")
lm += "Best answer: " + select(
["A) Paris", "B) London", "C) Berlin", "D) Madrid"],
name="answer"
)
print(lm["sentiment"])
print(lm["answer"])
3. Token Healing
Guidance automatically "heals" token boundaries between prompt and generation.
Problem: Tokenization creates unnatural boundaries.
prompt = "The capital of France is "
Solution: Guidance backs up one token and regenerates.
from guidance import models, gen
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
lm += "The capital of France is " + gen("capital", max_tokens=5)
Benefits:
- Natural text boundaries
- No awkward spacing issues
- Better model performance (sees natural token sequences)
4. Grammar-Based Generation
Define complex structures by composing grammar functions. The template-string
grammar= form is not part of current guidance — build grammars from
composable functions, or use guidance.json() for JSON.
from guidance import models, gen
from guidance import json as gen_json
from pydantic import BaseModel, Field
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
class Person(BaseModel):
name: str = Field(pattern=r"[A-Za-z ]+")
age: int
email: str = Field(pattern=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}")
lm += gen_json(name="person", schema=Person)
print(lm["person"])
grammar = "name=" + gen("name", regex=r"[A-Za-z ]+") + " age=" + gen("age", regex=r"[0-9]+")
lm += grammar
Use cases:
- Complex structured outputs
- Nested data structures
- Programming language syntax
- Domain-specific languages
5. Guidance Functions
Create reusable generation patterns with the @guidance decorator.
from guidance import guidance, gen, models
@guidance
def generate_person(lm):
"""Generate a person with name and age."""
lm += "Name: " + gen("name", max_tokens=20, stop="\n")
lm += "\nAge: " + gen("age", regex=r"[0-9]+", max_tokens=3)
return lm
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
lm = generate_person(lm)
print(lm["name"])
print(lm["age"])
Stateful Functions:
@guidance(stateless=False)
def react_agent(lm, question, tools, max_rounds=5):
"""ReAct agent with tool use."""
lm += f"Question: {question}\n\n"
for i in range(max_rounds):
lm += f"Thought {i+1}: " + gen("thought", stop="\n")
lm += "\nAction: " + select(list(tools.keys()), name="action")
tool_result = tools[lm["action"]]()
lm += f"\nObservation: {tool_result}\n\n"
lm += "Done? " + select(["Yes", "No"], name="done")
if lm["done"] == "Yes":
break
lm += "\nFinal Answer: " + gen("answer", max_tokens=100)
return lm
Backend Configuration
OpenAI (remote — unconstrained only)
Remote API backends cannot do constrained generation (regex/select/grammar);
use them only for plain chat/gen(). For constraints, use a local backend.
from guidance import models
lm = models.OpenAI(
model="gpt-4o-mini",
api_key="your-api-key"
)
Local Models (Transformers)
from guidance.models import Transformers
lm = Transformers(
"microsoft/Phi-4-mini-instruct",
device="cuda"
)
Local Models (llama.cpp)
from guidance.models import LlamaCpp
lm = LlamaCpp(
model_path="/path/to/model.gguf",
n_ctx=4096,
n_gpu_layers=35
)
Common Patterns
Pattern 1: JSON Generation
from guidance import models, gen, system, user, assistant
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
with system():
lm += "You generate valid JSON."
with user():
lm += "Generate a user profile with name, age, and email."
with assistant():
lm += """{
"name": """ + gen("name", regex=r'"[A-Za-z ]+"', max_tokens=30) + """,
"age": """ + gen("age", regex=r"[0-9]+", max_tokens=3) + """,
"email": """ + gen("email", regex=r'"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}"', max_tokens=50) + """
}"""
print(lm)
Pattern 2: Classification
from guidance import models, gen, select
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
text = "This product is amazing! I love it."
lm += f"Text: {text}\n"
lm += "Sentiment: " + select(["positive", "negative", "neutral"], name="sentiment")
lm += "\nConfidence: " + gen("confidence", regex=r"[0-9]+", max_tokens=3) + "%"
print(f"Sentiment: {lm['sentiment']}")
print(f"Confidence: {lm['confidence']}%")
Pattern 3: Multi-Step Reasoning
from guidance import models, gen, guidance
@guidance
def chain_of_thought(lm, question):
"""Generate answer with step-by-step reasoning."""
lm += f"Question: {question}\n\n"
for i in range(3):
lm += f"Step {i+1}: " + gen(f"step_{i+1}", stop="\n", max_tokens=100) + "\n"
lm += "\nTherefore, the answer is: " + gen("answer", max_tokens=50)
return lm
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
lm = chain_of_thought(lm, "What is 15% of 200?")
print(lm["answer"])
Pattern 4: ReAct Agent
from guidance import models, gen, select, guidance
@guidance(stateless=False)
def react_agent(lm, question):
"""ReAct agent with tool use."""
tools = {
"calculator": lambda expr: eval(expr),
"search": lambda query: f"Search results for: {query}",
}
lm += f"Question: {question}\n\n"
for round in range(5):
lm += f"Thought: " + gen("thought", stop="\n") + "\n"
lm += "Action: " + select(["calculator", "search", "answer"], name="action")
if lm["action"] == "answer":
lm += "\nFinal Answer: " + gen("answer", max_tokens=100)
break
lm += "\nAction Input: " + gen("action_input", stop="\n") + "\n"
if lm["action"] in tools:
result = tools[lm[]](lm[])
lm +=
lm
lm = models.Transformers()
lm = react_agent(lm, )
(lm[])
Pattern 5: Data Extraction
from guidance import models, gen, guidance
@guidance
def extract_entities(lm, text):
"""Extract structured entities from text."""
lm += f"Text: {text}\n\n"
lm += "Person: " + gen("person", stop="\n", max_tokens=30) + "\n"
lm += "Organization: " + gen("organization", stop="\n", max_tokens=30) + "\n"
lm += "Date: " + gen("date", regex=r"\d{4}-\d{2}-\d{2}", max_tokens=10) + "\n"
lm += "Location: " + gen("location", stop="\n", max_tokens=30) + "\n"
return lm
text = "Tim Cook announced at Apple Park on 2024-09-15 in Cupertino."
lm = models.Transformers("microsoft/Phi-4-mini-instruct")
lm = extract_entities(lm, text)
print(f"Person: {lm['person']}")
print(f"Organization: {lm['organization']}")
print(f"Date: {lm['date']}")
print()
Best Practices
1. Use Regex for Format Validation
lm += "Email: " + gen("email", regex=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}")
lm += "Email: " + gen("email", max_tokens=50)
2. Use select() for Fixed Categories
lm += "Status: " + select(["pending", "approved", "rejected"], name="status")
lm += "Status: " + gen("status", max_tokens=20)
3. Leverage Token Healing
lm += "The capital is " + gen("capital")
4. Use stop Sequences
lm += "Name: " + gen("name", stop="\n")
lm += "Name: " + gen("name", max_tokens=50)
5. Create Reusable Functions
@guidance
def generate_person(lm):
lm += "Name: " + gen("name", stop="\n")
lm += "\nAge: " + gen("age", regex=r"[0-9]+")
return lm
lm = generate_person(lm)
lm += "\n\n"
lm = generate_person(lm)
6. Balance Constraints
lm += gen("name", regex=r"[A-Za-z ]+", max_tokens=30)
lm += gen("name", regex=r"^(John|Jane)$", max_tokens=10)
Comparison to Alternatives
| Feature | Guidance | Instructor | Outlines | LMQL |
|---|
| Regex Constraints | ✅ Yes | ❌ No | ✅ Yes | ✅ Yes |
| Grammar Support | ✅ CFG | ❌ No | ✅ CFG | ✅ CFG |
| Pydantic Validation | ❌ No | ✅ Yes | ✅ Yes | ❌ No |
| Token Healing | ✅ Yes | ❌ No | ✅ Yes | ❌ No |
| Local Models | ✅ Yes | ⚠️ Limited | ✅ Yes | ✅ Yes |
| API Models | ✅ Yes | ✅ Yes | ⚠️ Limited | ✅ Yes |
| Pythonic Syntax | ✅ Yes | ✅ Yes | ✅ Yes | ❌ SQL-like |
| Learning Curve | Low | Low | Medium | High |
When to choose Guidance:
- Need regex/grammar constraints
- Want token healing
- Building complex workflows with control flow
- Using local models (Transformers, llama.cpp)
- Prefer Pythonic syntax
When to choose alternatives:
- Instructor: Need Pydantic validation with automatic retrying
- Outlines: Need JSON schema validation
- LMQL: Prefer declarative query syntax
Performance Characteristics
Latency Reduction:
- 30-50% faster than traditional prompting for constrained outputs
- Token healing reduces unnecessary regeneration
- Grammar constraints prevent invalid token generation
Memory Usage:
- Minimal overhead vs unconstrained generation
- Grammar compilation cached after first use
- Efficient token filtering at inference time
Token Efficiency:
- Prevents wasted tokens on invalid outputs
- No need for retry loops
- Direct path to valid outputs
Resources
See Also
references/constraints.md - Comprehensive regex and grammar patterns
references/backends.md - Backend-specific configuration
references/examples.md - Production-ready examples