Skip to main content
braintrustdata
GitHub-Creator-Profil

braintrustdata

Repository-Ansicht von 63 gesammelten Skills in 14 GitHub-Repositories.

gesammelte Skills
63
Repositories
14
aktualisiert
27. Aug. 2026
Hier werden die Top 8 Repositories angezeigt; die vollständige Repository-Liste folgt darunter.
Repository-Explorer

Repositories und repräsentative Skills

braintrust-analyze-eval-experiment
nicht klassifiziert

Analyze completed LLM or agent eval experiments using uncertainty-aware and decision-relevant methods. Use to audit run completeness and pairing, calculate confidence intervals, run paired comparisons, report wins, losses, and ties, incorporate run-to-run…

17. Aug. 2026
braintrust-attribute-multi-variable-change
nicht klassifiziert

Attribute an observed change when several things moved at once — model plus prompt plus tools, a provider migration, a framework upgrade, or a vendor swap that bundles serving stack with model. Use when asked which part of a change caused the result, when a…

17. Aug. 2026
braintrust-build-eval-dataset
nicht klassifiziert

Create, edit, audit, or compare eval datasets for LLM applications and agents, including target-population definition, case sourcing from production traces, stratified sampling, label provenance and label audits, expected values as constraints for open-ended…

17. Aug. 2026
braintrust-define-eval-objective
nicht klassifiziert

Create, edit, or audit an eval objective by working backward from a product decision to the target outcome, construct, population, intended claim, and verification-versus-validation questions. Use when a team is unsure what an eval should establish, asks…

17. Aug. 2026
braintrust-define-eval-release-gate
nicht klassifiziert

Create, edit, audit, or apply release gates for LLM applications and agents. Use to combine minimum meaningful improvement, statistical significance, regression rate, subgroup consistency, worst-run stability, all-attempts reliability, safety upper bounds,…

17. Aug. 2026
braintrust-deploy-evaluator
nicht klassifiziert

Take a validated scorer or classifier from definition to running instrument in Braintrust — scope selection, inline testing before saving, saving as an evaluator, attaching an online-scoring rule, activating it for new traffic, and backfilling history with a…

17. Aug. 2026
braintrust-design-eval-experiment
nicht klassifiziert

Design or audit controlled eval experiments for model, prompt, retrieval, tool, guardrail, or agent-architecture changes. Use before data collection to state directional and minimum-effect hypotheses, name independent, dependent, and control variables…

17. Aug. 2026
braintrust-design-eval-instrumentation
nicht klassifiziert

Design the trace and eval-dataset schema for an LLM app or agent, and wire the system to emit it. Use when deciding what to log, designing a trace schema, setting up tracing or observability before evals, or when failures cannot be debugged or sliced from…

17. Aug. 2026
Es werden 8 von 24 gesammelten Skills angezeigt.
troubleshoot-braintrust-mcp
nicht klassifiziert

This plugin auto-configures a "braintrust" MCP server. If you can't see it or reach it, activate this skill

27. Aug. 2026
add-coding-agent-capture
Softwareentwickler

Add or review live coding-agent event capture through blocking command hooks or an in-process plugin such as a JavaScript adapter. Use when forwarding native lifecycle, model, tool, permission, or subagent events into the Braintrust daemon, or when fixing…

10. Aug. 2026
add-coding-agent-import
Softwareentwickler

Add or review historical transcript import and live attach support for a coding agent, including session lookup, native parsing, synthetic lifecycle envelopes, shared translator use, incremental tailing, destination overrides, and import tests. Use when…

10. Aug. 2026
add-coding-agent-integration
Softwareentwickler

Orchestrate a complete Braintrust coding-agent integration by delegating feasibility, daemon translation, event capture, setup, managed run, transcript import, verification, and shipping to the specialized repo-local skills. Use for end-to-end support for a…

10. Aug. 2026
add-coding-agent-run
Softwareentwickler

Add or review invocation-local managed-run support for a coding agent, including executable dispatch, temporary hook or plugin injection, route isolation, duplicate-capture suppression, process behavior, trust safety, and run-command tests. Use when…

10. Aug. 2026
add-coding-agent-setup
Softwareentwickler

Add or review persistent setup for a coding-agent tracing integration, including public CLI exposure, plugin or hook installation, non-secret route configuration, idempotent updates, disablement, and isolated setup tests. Use when implementing or fixing the…

10. Aug. 2026
add-coding-agent-translator
Softwareentwickler

Implement or review a coding agent's daemon translator, including source registration, native event correlation, trace shape, deterministic recovery, routing metadata, and translator tests. Use when adding a new agent translator or changing how an existing…

10. Aug. 2026
ship-coding-agent-integration
Softwareentwickler

Package and release a coding-agent tracing integration through its marketplace, package registry, or distribution repository, including manifests, builds, validation, versioning, CI, documentation, publishing automation, and installed-artifact smoke tests.…

10. Aug. 2026
Es werden 8 von 12 gesammelten Skills angezeigt.
sdk-integrations
Softwareentwickler

Create or update Braintrust Python SDK integrations built on the integrations API under `py/src/braintrust/integrations/`. Use when adding a new integration package, extending an existing provider integration, changing patchers, tracing, manual `wrap_*()`…

30. Juli 2026
sdk-dependency-updates
Softwareentwickler

Review and refresh Braintrust Python SDK dependency update PRs, especially automated `chore(deps): weekly dependency update` PRs. Use when Codex needs to inspect `py/pyproject.toml` and `py/uv.lock`, reproduce the workflow's label decision, decide whether…

1. Mai 2026
sdk-benchmarking
Softwareentwickler

Run, compare, and extend Braintrust Python SDK pyperf benchmarks. Use when touching hot-path code in `py/src/braintrust/` such as serialization, deep-copy, span creation, or logging; when adding or updating files under `py/benchmarks/`; or when you need…

15. Apr. 2026
sdk-ci-triage
Softwareentwickler

Triage and reproduce Braintrust Python SDK CI failures. Use when asked why CI failed, to fix broken CI on a PR, to inspect a failing GitHub Actions job, or to map a failing matrix job back to the exact local nox session, provider version, or workflow step…

15. Apr. 2026
sdk-vcr-workflows
Softwarequalitätssicherungsanalysten und -tester

Work with Braintrust Python SDK VCR and cassette-backed tests. Use when adding or updating cassette-backed provider tests, deciding whether to re-record cassettes, debugging VCR failures, converting mock-heavy coverage to VCR, or handling cassette hygiene for…

15. Apr. 2026
commit-message
Softwareentwickler

Suggest a Braintrust SDK repo-style commit message from the current diff and conversation. Use when asked to write, suggest, or generate a commit message for the current changes.

13. Apr. 2026
12 von 14 Repositories angezeigt