Skip to main content
togethercomputer
ملف منشئ GitHub

togethercomputer

عرض على مستوى المستودعات لـ ٢٨ skills مجمعة عبر ٤ مستودعات GitHub.

skills مجمعة
٢٨
مستودعات
٤
محدث
١٧ أغسطس ٢٠٢٦
مستكشف المستودعات

المستودعات و skills الممثلة

together-gpu-clusters
مديرو الشبكات وأنظمة الحاسوب

On-demand and reserved GPU clusters (H100, H200, B200) on Together AI with Kubernetes or Slurm orchestration, shared storage, credential management, and cluster scaling for ML and HPC jobs. Reach for it when the user needs multi-node compute or infrastructure…

١٧ أغسطس ٢٠٢٦
together-audio
مطوّرو البرمجيات

Text-to-speech and speech-to-text via Together AI, including REST, streaming, and realtime WebSocket TTS, plus transcription, translation, diarization, timestamps, and live STT. Reach for it whenever the user needs audio in or audio out on Together AI rather…

٧ أغسطس ٢٠٢٦
together-chat-completions
مطوّرو البرمجيات

Real-time and streaming text generation via Together AI's OpenAI-compatible chat/completions API, including multi-turn conversations, tool and function calling, structured JSON outputs, and reasoning models. Reach for it whenever the user wants to build or…

٢٣ يوليو ٢٠٢٦
together-evaluations
مطوّرو البرمجيات

LLM-as-a-judge evaluation framework on Together AI. Classify, score, and compare model outputs, select judge models, use external-provider judges or targets, poll results and download reports. Reach for it whenever the user wants to benchmark outputs, grade…

٢٣ يوليو ٢٠٢٦
together-dedicated-containers
مطوّرو البرمجيات

Custom Dockerized inference workers on Together AI's managed GPU infrastructure. Build with Sprocket SDK, configure with Jig CLI, submit async queue jobs, and poll results. Reach for it whenever the user needs container-level control rather than a standard…

٢٣ يوليو ٢٠٢٦
together-dedicated-model-inference
مطوّرو البرمجيات

Deploy and operate models on dedicated GPUs with Together AI's Dedicated Model Inference (DMI, the v2 dedicated endpoints API): beta endpoints, deployments, deployment profiles and hardware configs, autoscaling, traffic splitting, A/B tests, shadow…

٢٣ يوليو ٢٠٢٦
together-embeddings
مطوّرو البرمجيات

Dense vector embeddings, semantic search, RAG pipelines, and reranking via Together AI. Generate embeddings with open-source models and rerank results behind dedicated endpoints. Reach for it whenever the user needs vector representations or retrieval quality…

٢٣ يوليو ٢٠٢٦
together-fine-tuning
مطوّرو البرمجيات

LoRA, full fine-tuning, DPO preference tuning, VLM training, function-calling tuning, reasoning tuning, and BYOM uploads on Together AI. Reach for it whenever the user wants to adapt a model on custom data rather than only run inference, evaluate outputs, or…

٢٣ يوليو ٢٠٢٦
عرض 8 من أصل ١٤ skills مجمعة.
add-jit-kernel
غير مصنف

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jit_kernel module

٩ أغسطس ٢٠٢٦
cookbook-add-model
غير مصنف

Add a new model to the SGLang Cookbook (docs/, Mintlify), config-driven format — instantiate the model-agnostic template into a per-model config (+ benchmarks) JSX under src/snippets/configs/, an MDX page, the docs.json nav entry, NEW-tag hygiene, and the…

٩ أغسطس ٢٠٢٦
llm-torch-profiler-analysis
غير مصنف

Unified LLM torch-profiler triage skill for `sglang`, `vllm`, `TensorRT-LLM`, and `TokenSpeed`. Use it to inspect an existing `trace.json(.gz)` or profile directory, or to drive live profiling against a running server when supported and return one three-table…

٩ أغسطس ٢٠٢٦
sglang-runtime-context
غير مصنف

How SGLang's runtime configuration and process-global state are organized (RuntimeContext tiers, publish + namespace config bags, the pristine ServerArgs seed, override entry points, resource/stream/buffer leases, per-forward flags), the CI guardrails that…

٩ أغسطس ٢٠٢٦
sglang-diffusion-add-model
غير مصنف

Use when adding a new diffusion model or Diffusers pipeline to SGLang.

٩ أغسطس ٢٠٢٦
sglang-diffusion-benchmark-profile
غير مصنف

Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

٩ أغسطس ٢٠٢٦
sglang-diffusion-modelopt-quant
غير مصنف

Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

٩ أغسطس ٢٠٢٦
sglang-diffusion-performance
غير مصنف

Use when choosing the fastest SGLang Diffusion flags for a model, GPU, and VRAM budget.

٩ أغسطس ٢٠٢٦
عرض ٤ من أصل ٤ مستودعات
تم تحميل كل المستودعات