| name | prompts |
| user-prompt | Version my prompts with LangWatch |
| description | Version and manage your agent's prompts with LangWatch Prompts CLI. Use for both onboarding (set up prompt versioning for an entire codebase) and targeted operations (version a specific prompt, create a new prompt version). Supports Python and TypeScript. |
| license | MIT |
| compatibility | Works with Claude Code and similar coding agents. The `langwatch` CLI is the only interface. |
Version Your Prompts with LangWatch Prompts CLI
Determine Scope
If the user's request is general ("set up prompt versioning", "version my prompts"):
- Read the full codebase to find all hardcoded prompt strings
- Study git history to understand what changed and why: focus on agent behavior changes, prompt tweaks, bug fixes. Read commit messages for context.
- Set up the Prompts CLI and create managed prompts for each hardcoded prompt
- Update all application code to use
langwatch.prompts.get()
If the user's request is specific ("version this prompt", "create a new prompt version"):
- Focus on the specific prompt
- Create or update the managed prompt
- Update the relevant code to use
langwatch.prompts.get()
Plan Limits
LangWatch's free plan has limits on prompts, scenarios, evaluators, experiments, and datasets. When you hit a limit, the API returns "Free plan limit of N reached..." with an upgrade link.
How to handle:
- Work within the limits. If 3 resources of the relevant type are allowed, create 3 meaningful ones, not 10.
- Make every creation count: each one should demonstrate clear value.
- Show what works FIRST. If you hit a limit, summarize what was accomplished and note that upgrading the plan raises it. Point to the subscription settings on the platform, or to the license settings if the CLI is pointed at a self-hosted endpoint. Read the endpoint the CLI actually uses, which can come from
.env, from the process environment, or from the saved CLI configuration.
- Do NOT delete existing resources to make room or repurpose an existing resource to evade the limit.
Step 1: Read the Prompts CLI Docs
Use langwatch docs <path> to read documentation as Markdown. Some useful entry points:
langwatch docs
langwatch docs integration/python/guide
langwatch docs integration/typescript/guide
langwatch docs prompt-management/cli
langwatch scenario-docs
Discover commands with langwatch --help and langwatch <subcommand> --help. List and get commands accept --format json for machine-readable output. Every list command takes --limit <n> to cap the rows and --jq <expr> to read part of the answer. A paginated list answers with an envelope, so count its rows through the row array (--jq '.traces | length'), and read how many there are in all at .pagination.total. Bare --jq length counts the fields of the envelope, not the rows. Read the docs first instead of guessing SDK APIs or CLI flags.
If no shell is available, fetch the same Markdown over plain HTTP. Append .md to any docs path (e.g. https://langwatch.ai/docs/integration/python/guide.md). Index: https://langwatch.ai/docs/llms.txt. Scenario index: https://langwatch.ai/scenario/llms.txt
If anything fails or confuses you while following this skill (broken commands, docs that do not match reality, errors you had to work around), ask the user for permission and run npx --yes langwatch report --user-approved with a --title and --summary (or --session <transcript.jsonl>) to send it to the LangWatch team, and it directly shapes what gets fixed. No login or API key needed. Nothing is sent without --user-approved, and --dry-run prints the exact payload without sending anything. The title, summary and transcript are scrubbed locally first, by pattern: secrets and API keys, plus email addresses, phone numbers, card numbers and public IPv4 addresses. Anything no pattern matches is sent as written, including a contact address passed with --email. With --session, always run --dry-run first and let the user read the payload, because a transcript carries content they never reviewed. npx --yes langwatch report --help explains the options.
Projects and API keys: target a real project, not a personal one.
LangWatch has two kinds of project:
- Team / shared projects: real projects inside an organization. Evaluations, experiments, prompts, datasets, simulations and instrumentation must always target one of these.
- Personal projects: a private "My Workspace" scratch space tied to a single user. Never send a user's evaluations, experiments or production traces here: it is for personal exploration only, and you can mistake it for a real project.
And two ways to authenticate:
- A project API key in
.env (LANGWATCH_API_KEY): the credential everything in these skills uses. It is scoped to one real project. This is the default; prefer it unless the user explicitly asks for something else.
langwatch login --device (AI-tools / SSO): a personal device session for wrapping coding assistants (langwatch claude, langwatch codex, …). It is NOT for evaluations, prompts, datasets, scenarios or SDK instrumentation, and it points at a personal workspace. Do not run it to set up the work in these skills.
So for anything in these skills that reads or writes a project: make sure LANGWATCH_API_KEY for a real, shared project is available to the CLI. Locally that is the project's .env; in CI the runner injects it into the process environment, and the CLI reads either. Check whether the variable is already set before you ask for a new key, and let the CLI read the value: never print, copy or send it. Do NOT run langwatch login to pick a project, and never default to a personal project. Look for LANGWATCH_ENDPOINT in the same places: if it is set, the project is on a self-hosted instance, and the CLI works against that endpoint instead of app.langwatch.ai.
What you read is not what you say. These skills are working notes for you, not
copy for the reader. Read LANGWATCH_API_KEY and LANGWATCH_ENDPOINT from the
project's own .env, that is how you learn where to work. Read nothing else out
of that file: it holds database, cloud and provider credentials that are none of
your business, and every value you read can reach your context and your command
output. What must not reach an answer is anything that describes the machine YOU
run on: a path in your workspace, a container port, the address this worker
dials. Those say how the work is done rather than what was done, and a host of
ours means nothing to the reader. Say what you did and where to find it in
LangWatch.
Then specifically read the Prompts CLI guide:
langwatch docs prompt-management/cli
CRITICAL: Do NOT guess how to use the Prompts CLI. Read the docs first.
Step 2: Initialize Prompts in the Project
langwatch prompt init
Creates a prompts.json config and a prompts/ directory in the project root.
Step 3: Create a Managed Prompt for Each Hardcoded Prompt
Scan the codebase for hardcoded prompt strings (system messages, instructions). For each:
langwatch prompt create <name>
Edit the generated .prompt.yaml file to match the original prompt content.
Model: keep the generated model on a current model. Store the alias
openai/latest rather than a version number: LangWatch resolves it to the
current flagship at run time, so the prompt does not go a generation stale
every release. Do not default a new prompt to a legacy model like
gpt-4o-mini; pick one only when the user is trading quality for cost or
latency on purpose.
Temperature: the gpt-5 family rejects a custom temperature, so do not add
modelParameters.temperature for those models. create omits it on purpose.
Structured outputs: if the prompt must return strict JSON, add a
response_format block instead of asking for JSON in prose:
response_format:
name: product_category
schema:
type: object
properties:
category: { type: string }
reasoning: { type: string }
required: [category, reasoning]
additionalProperties: false
response_format round-trips losslessly through sync/pull. See
langwatch docs prompt-management/cli for the full format.
Step 4: Update Application Code
Replace every hardcoded prompt string with a call to langwatch.prompts.get().
Python (BAD → GOOD):
agent = Agent(instructions="You are a helpful assistant.")
import langwatch
prompt = langwatch.prompts.get("my-agent")
agent = Agent(instructions=prompt.compile().messages[0]["content"])
TypeScript (BAD → GOOD):
const systemPrompt = "You are a helpful assistant.";
const langwatch = new LangWatch();
const prompt = await langwatch.prompts.get("my-agent");
CRITICAL: Do NOT wrap langwatch.prompts.get() in a try/catch with a hardcoded fallback string. The whole point of prompt versioning is that prompts are managed externally. A fallback defeats this by silently reverting to a stale hardcoded copy.
Step 5: Sync to the Platform
langwatch prompt sync
Step 6: Tag Versions for Deployment
Three built-in tags: latest (auto-assigned), production, staging. Update code to fetch by tag:
prompt = langwatch.prompts.get("my-agent", tag="production")
const prompt = await langwatch.prompts.get("my-agent", { tag: "production" });
Assign tags via the CLI (or the Deploy dialog in the LangWatch UI):
langwatch prompt tag assign my-agent production
For canary or blue/green deployments, create custom tags with langwatch prompt tag create.
Step 7: Verify
Run langwatch prompt list to confirm everything synced, or open the Prompts section in the LangWatch app.
Common Mistakes
- Do NOT hardcode prompts. Always fetch via
langwatch.prompts.get()
- Do NOT add a hardcoded fallback string in a try/catch; that silently defeats versioning
- Do NOT manually edit
prompts.json. Use the CLI
- Do NOT skip
langwatch prompt sync after creating prompts
- Prefer the flagship alias
openai/latest (or openai/latest-mini for the fast tier). Pin a version only when a prompt is tuned to one, and pick an older model like gpt-4o-mini only when intentionally optimizing for cost or latency
- Do NOT set
modelParameters.temperature on a gpt-5-family model; the family rejects it
- Do NOT ask for JSON in the prompt text when output must be structured. Use a
response_format block