Skip to main content
langwatch
ملف منشئ GitHub

langwatch

عرض على مستوى المستودعات لـ ٥٤ skills مجمعة عبر ٤ مستودعات GitHub.

skills مجمعة
٥٤
مستودعات
٤
محدث
٥ سبتمبر ٢٠٢٦
مستكشف المستودعات

المستودعات و skills الممثلة

connect-agent
محللو ضمان جودة البرمجيات والمختبرون

Connect the codebase's AI agent to LangWatch agent simulations, so test suites run against the real agent process. Adds a small connect function beside the service startup that calls the agent already in the codebase, which opens an outbound connection and…

٥ سبتمبر ٢٠٢٦
scenarios
محللو ضمان جودة البرمجيات والمختبرون

Test your AI agent with simulation-based scenarios. Covers writing scenario test code (Scenario SDK), creating platform scenarios via the `langwatch` CLI against a connected agent, reading the run parameters that agent declares so the scenarios and comparison…

١ سبتمبر ٢٠٢٦
agent-best-practices
مطوّرو البرمجيات

Expert AI engineering consultant for your agent development practices. Audits your codebase, traces, evaluations, and scenarios against best practices, then guides you to close the gaps, starting from low-hanging fruit and going deeper. Use when you want to…

٢٨ أغسطس ٢٠٢٦
prompt-optimization
مطوّرو البرمجيات

Improve a prompt on the evaluations workbench through a measured loop. Score the baseline first, then duplicate the target column, form a hypothesis from failing rows, edit the copy's prompt draft, run, compare pass rate and cost, and repeat until the numbers…

٢٧ أغسطس ٢٠٢٦
context-sweet-spot
مطوّرو البرمجيات

Investigates the context economics of your own coding-agent sessions in LangWatch. Reads real sessions to find where carrying a fat context stops paying for itself, measured in cache rebuilds, compactions and cost per turn, and delivers a report with the…

٢٧ أغسطس ٢٠٢٦
provider-cost-comparison
محللو التمويل والاستثمار

Prices your real LangWatch usage mix against other model providers. Exports your actual token mix per model, including the cache read and write split, fetches current price cards, and reprices the same month of usage under each candidate, with the cache…

٢٧ أغسطس ٢٠٢٦
agent-improve
مطوّرو البرمجيات

Turns production evidence into tested improvements for your AI agent. Forms hypotheses from real traces and analytics, explains the reasoning behind each one, then executes with the user: scenario tests that reproduce production failures, prompt and code…

٢٦ أغسطس ٢٠٢٦
drive-the-ui
مطوّرو البرمجيات

Drive the page the user has open through live UI actions. List the actions a page accepts, call them with typed payloads, and read the live state including unsaved edits. Use when the user is looking at a page you can operate, such as the evaluations…

٢٦ أغسطس ٢٠٢٦
عرض 8 من أصل ٣٣ skills مجمعة.
context-sweet-spot
غير مصنف

Investigates the context economics of your own coding-agent sessions in LangWatch. Reads real sessions to find where carrying a fat context stops paying for itself, measured in cache rebuilds, compactions and cost per turn, and delivers a report with the…

٢٧ أغسطس ٢٠٢٦
provider-cost-comparison
غير مصنف

Prices your real LangWatch usage mix against other model providers. Exports your actual token mix per model, including the cache read and write split, fetches current price cards, and reprices the same month of usage under each candidate, with the cache…

٢٧ أغسطس ٢٠٢٦
agent-improve
غير مصنف

Turns production evidence into tested improvements for your AI agent. Forms hypotheses from real traces and analytics, explains the reasoning behind each one, then executes with the user: scenario tests that reproduce production failures, prompt and code…

٢٦ أغسطس ٢٠٢٦
agent-performance
غير مصنف

Deep-dive diagnosis of how your AI agent behaves in production. Explores LangWatch analytics and traces end to end to map failure patterns, dissatisfied users, token cost hotspots, edge cases, behavior changes, and outliers, then delivers an HTML report where…

٢٦ أغسطس ٢٠٢٦
connect-agent
غير مصنف

Connect the codebase's AI agent to LangWatch agent simulations over HTTP, so scenario suites run against it from the platform. Finds or adds the agent's chat endpoint, wires authentication for scenario traffic, makes the server adopt the W3C traceparent…

٢٦ أغسطس ٢٠٢٦
experiments
غير مصنف

Create and run LangWatch experiments for pre-deployment batch testing. Use when the user wants to test an agent against a dataset, compare prompts or models, benchmark quality, detect regressions, or add a CI quality gate. Do not use for production monitoring…

٢٦ أغسطس ٢٠٢٦
level-up
غير مصنف

Take your AI agent to the next level with full LangWatch integration. Adds tracing, prompt versioning, evaluation experiments, and simulation tests in one go. Use when the user wants comprehensive observability, testing, and prompt management for their agent.

٢٦ أغسطس ٢٠٢٦
online-evaluations
غير مصنف

Configure LangWatch online evaluations and guardrails for production traffic. Use when the user wants to score live traces or threads, monitor production quality, sample incoming traffic, or synchronously block unsafe requests and responses. Do not use for…

٢٦ أغسطس ٢٠٢٦
عرض 8 من أصل ١٧ skills مجمعة.
عرض ٤ من أصل ٤ مستودعات
تم تحميل كل المستودعات