Skip to main content

이 저장소의 skills

NVIDIA-TAO/tao-skill-bank - 2페이지

SkillsMP는 NVIDIA-TAO/tao-skill-bank에서 81개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.

NVIDIA-TAO/tao-skill-bank

수집된 skill 81개 중 40개를 표시합니다.

직업 분류
소프트웨어 개발자
설명

BEVFusion for multi-sensor 3D object detection. Fuses LiDAR point clouds and camera images in bird's-eye-view (BEV) space, used in autonomous driving for robust 3D perception. Use when training, evaluating, or running inference for a TAO BEVFusion model.…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

CenterPose for keypoint / pose estimation. Detects object centers and regresses keypoint locations for 6-DoF object pose estimation. Use when training, evaluating, exporting, or running inference for a TAO CenterPose model. Trigger phrases include "train…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Deformable DETR for 2D object detection. Uses deformable attention for efficient multi-scale feature processing, lighter than DINO with competitive accuracy. Use when training, evaluating, exporting, quantizing, or running inference for a TAO Deformable-DETR…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures. Predicts per-pixel depth from single RGB images. Use when training, evaluating, exporting, or running inference for a TAO monocular depth model. Trigger…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support. Use when training, evaluating, exporting, distilling, quantizing, or running…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Stereo depth estimation using FoundationStereo. Predicts disparity maps from stereo image pairs for 3D reconstruction. Use when training, evaluating, exporting, or running inference for a TAO FoundationStereo model. Trigger phrases include "train stereo…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Grounding DINO for open-set object detection. Combines DINO-style detection with a BERT text encoder for language-guided detection — detects objects described by text prompts without a fixed class vocabulary. Use when training, evaluating, exporting,…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Masked Auto-Encoder (MAE) for self-supervised pretraining and fine-tuning. Masks random patches and reconstructs them to learn visual representations; supports pretrain and finetune stages. Use when training, evaluating, exporting, or running inference for a…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

MAL (Mask Auto-Label) for weakly-supervised segmentation. Produces segmentation masks from minimal annotations (point or box annotations) using a ViT-MAE backbone. Use when training, evaluating, or running inference for a TAO MAL model. Trigger phrases…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Mask Grounding DINO for grounded instance segmentation. Extends Grounding DINO with a mask-prediction head for open-set segmentation guided by text prompts. Use when training, evaluating, exporting, quantizing, or running inference for a TAO…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Mask2Former for universal image segmentation (panoptic, instance, and semantic). Transformer-based with masked attention for high-quality segmentation results. Use when training, evaluating, exporting, quantizing, or running inference for a TAO Mask2Former…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Metric-learning recognition (ml-recog) for fine-grained visual recognition. Learns embeddings for retrieval-based matching (e.g., retail product recognition) using triplet / contrastive losses. Use when training, evaluating, exporting, or running inference…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

NVDINOv2 for self-supervised visual representation learning. Trains vision transformers via self-distillation (teacher-student) without labels and produces general-purpose visual features. Use when training, exporting, or running inference for a TAO NVDINOv2…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

NVPanoptix3D for panoptic 3D scene reconstruction from posed RGB images. Produces 3D panoptic segmentation (semantic, instance, and panoptic masks) with occupancy completion. Built on a VGGT backbone with a Mask2Former-style head and 3D frustum…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

OCDNet for scene text detection. Detects arbitrary-oriented text regions in natural images using a differentiable binarization approach. Use when training, evaluating, exporting, pruning, quantizing, retraining, or running inference for a TAO OCDNet model.…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

OCRNet for scene text recognition. Recognizes text content from cropped text-region images and supports CTC and attention-based decoders. Use when training, evaluating, exporting, pruning, quantizing, retraining, or running inference for a TAO OCRNet model.…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

OneFormer for universal image segmentation. Unifies panoptic, instance, and semantic segmentation with a single architecture using task-conditioned queries. Use when training, evaluating, exporting, quantizing, or running inference for a TAO OneFormer model.…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Optical Inspection for defect detection using Siamese networks. Compares image pairs to detect manufacturing defects, anomalies, or quality issues. Use when training, evaluating, exporting, or running inference for a TAO Optical Inspection model on AOI /…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

PointPillars for 3D object detection from LiDAR point clouds. Encodes point clouds into a pseudo-image via a pillar-based representation, then applies 2D detection — used in autonomous driving and robotics. Use when training, evaluating, exporting, pruning,…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Pose classification using ST-GCN (Spatial Temporal Graph Convolutional Network). Classifies skeleton sequences into action categories from pose-keypoint data. Use when training, evaluating, exporting, or running inference for a TAO pose-classification model.…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Person re-identification (ReID). Learns discriminative embeddings to match the same person across different camera views, based on metric learning. Use when training, evaluating, exporting, or running inference for a TAO person re-identification model.…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

RT-DETR (Real-Time DEtection TRansformer) for 2D object detection. Designed for real-time inference with competitive accuracy and supports distillation and quantization for deployment optimization. Use when training, evaluating, distilling, quantizing,…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

SegFormer for semantic segmentation. Lightweight transformer-based architecture with hierarchical feature extraction, efficient for real-time segmentation tasks. Use when training, evaluating, exporting, quantizing, or running inference for a TAO SegFormer…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Run a TAO training/evaluation/inference container on an NVIDIA Brev GPU instance. Instance provisioning (create/search/stop/delete/login) is delegated to the official brev-cli agent skill or the Brev MCP server; this skill covers only the TAO-specific part —…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Port a published computer vision paper's official code and training recipe onto a customer's own dataset, or diagnose why such a transfer produced bad numbers. Use this whenever someone wants to reproduce a CV paper, run a paper's repo on their own images,…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Runs the DEFT embed-then-mine workflow for VCN AOI iterations — embeds the gap-analysis target parquet, embeds a source pool, and mines nearest-neighbour source images for downstream augmentation. Use as the immediate next step after…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Standard single-step train/eval/export workflow for any TAO model. Use when training a TAO model on a dataset without iterative data augmentation, AutoML, or DEFT loops. Trigger phrases include "single train run", "train then evaluate then export", "plain TAO…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

The contract home for TAO's SDK-free execution pipeline — authoritative JSON Schemas for the four typed artifacts (spec-bundle, job-record, results_dir layout, best_rec) plus the fixed job-status vocabulary and the nested-not-dotted spec rule. Use when…

원문 언어: 영어

업데이트
직업 분류
기타 컴퓨터 관련 직업
설명

Answer what the TAO Skill Bank plugin can do by generating the response from packaged application, data, model, AutoML, and platform manifests. Use when the user asks "what can TAO Skill Bank do", "list TAO models", "which TAO workflows are available", or…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Run `tao-daft convert` to convert NVIDIA TAO DAFT datasets between supported formats. Do not use for non-DAFT data. Use when the user asks to convert a DAFT dataset, change DAFT format, change a TAO dataset format, or run `tao-daft convert`.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Run `tao-daft validate` to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors. Do not use for non-DAFT formats. Use when the user asks to validate a DAFT dataset, check DAFT schema, validate a TAO dataset format, or run `tao-daft…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Sparse4D for multi-camera temporal 3D object detection and tracking. Uses sparse queries with deformable attention across camera views and time for end-to-end 3D perception, with an instance bank for temporal tracking. Use when training, evaluating,…

원문 언어: 영어

업데이트
직업 분류
네트워크·컴퓨터 시스템 관리자
설명

Where and how GPU jobs run on this platform. One-to-three-sentence summary. Use when the user asks to "deploy on REPLACE-PLATFORM", "run on REPLACE-PLATFORM", or mentions the platform's distinctive concepts (e.g., resource shape, instance, node group).

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation. Use when analyzing ChangeNet model failures, investigating poor recall / FAR / PASS-NO_PASS metrics, auditing visual…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning traces, via VLM/LLM distillation. Use when the user wants to…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Routes the weakest VCN samples (output of `tao-analyze-gaps-visual-changenet`) into per-augmentation-module subsets based on each module's label eligibility. Use when the user asks to "route VCN gap samples", "split AOI gaps for k-NN mining and AnomalyGen",…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

DINOv3 continual self-supervised pre-training. Domain-adapts public DINOv3 ViT backbones on unlabeled images via teacher-student self-distillation (DINO + iBOT + KoLeo, optional Gram anchoring) and converts the EMA teacher into a timm-format backbone for…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Real-time stereo depth estimation using FastFoundationStereo (FFS), the distilled bp2 commercial variant of FoundationStereo. Predicts disparity maps from stereo image pairs with ~10× lower latency than full FoundationStereo. Use when training, evaluating,…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

One-to-three-sentence description of what this data transformation does. Use when the user asks to "REPLACE-WITH-INTENT", or mentions REPLACE-WITH-DOMAIN-TERMS. Include literal trigger phrases the user is likely to say.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

One-to-three-sentence description of what the skill does and when to use it. Use when the user asks to "REPLACE-WITH-INTENT-1", "REPLACE-WITH-INTENT-2", or mentions REPLACE-WITH-DOMAIN-TERMS. Include literal trigger phrases the user is likely to say.

원문 언어: 영어

업데이트
수집된 skill 81개 중 40개를 표시합니다.