Skip to main content

xinyangli/nixos-config

SkillsMP 已收集 xinyangli/nixos-config 中的 21 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
21
GitHub 星标
4
GitHub Forks
0

这个仓库中的 skills

职业分类待补全

已展示 21 / 21 个已收集 Skill。

职业分类
未分类
描述

Recover and retry failed records from Gemini batch annotation jobs. Covers grid-level retry (no re-concat), full re-concat retry pipelines, and after concat — how to upload manifest to S3, split shard ranges across teammates, run with --use-concat,…

原文语言:英语

更新
职业分类
未分类
描述

Update Canva ML model weights through W&B registry upload, artifact mirroring, model lockfile PRs, and staged deployment PRs. Use when updating ingredient-generation/media-transformation model checkpoints, running arnold registry upload, following Canva…

原文语言:英语

更新
职业分类
未分类
描述

Package image datasets into bucketed Parquet shards with resolution and aspect ratio bucketing. Use when converting JSONL+image datasets to Parquet format, creating training-ready datasets, or when the user mentions parquet, packaging, bucketing, sharding, or…

原文语言:英语

更新
职业分类
未分类
描述

Run verifier gates for Core CN dataset processing stages: S3 counts, parquet schema, row alignment, metadata naming, bbox validity, OCR/caption validity, resized image audits, and final sidecar quality.

原文语言:英语

更新
职业分类
未分类
描述

Orchestrate Core CN image dataset processing from raw assets through CPU filtering, dedup, main/sidecar Parquet packaging, 512/1024 resized Parquet, layer detection, HunyuanOCR, JSON captioning, score enrichment, and verifier gates. Use when planning or…

原文语言:英语

更新
职业分类
未分类
描述

Build 2x2 grid concat images from S3 image datasets for batch Gemini captioning. Use when creating concat grids, running the annotation batch infer pipeline, generating concat manifests, managing the prepare→concat→worker→run→postprocess workflow, or…

原文语言:英语

更新
职业分类
未分类
描述

Build 512-area and 1024-area resized image Parquet datasets, preserve row keys, audit valid/invalid rows, and feed crop metadata back into structured_description bbox columns.

原文语言:英语

更新
职业分类
未分类
描述

Reduce YubiKey touch frequency on Canva Linux devboxes. Diagnoses where touches come from (tsh/kubectl, Python kubernetes clients like utp, AWS prod profile) and applies workarounds. Use when the user says "yubikey 触发太频繁", "kubectl 每次要 touch", "aws cli 一直要…

原文语言:英语

更新
职业分类
未分类
描述

Start an HTTP service in Devbox and expose it to colleagues via infra highway tunnel. Covers tmux session setup, python http.server, highway tunnel creation, and warp access. Use when needing to share HTML demos, comparison pages, Streamlit/Gradio apps, or…

更新
职业分类
未分类
描述

Build image editing dataset pipelines using third-party APIs (OpenAI gpt-image-1.5/bluefire-alpha, Google Gemini). Covers generating edit instructions via Vertex AI batch, executing image edits via OpenAI Batch API or local concurrent API calls,…

原文语言:英语

更新
职业分类
未分类
描述

Multi-model quality assessment and data cleaning pipeline for image editing datasets. Evaluates (source, instruction, edited) triplets on three dimensions (edit adherence, visual quality, content preservation) using Gemini batch + OpenAI online judges.…

原文语言:英语

更新
职业分类
未分类
描述

Run layer detection, identify text-bearing elements, hand layer detection plus HunyuanOCR outputs to JSON caption jobs, and validate structured_description bbox outputs for final sidecar merge.

原文语言:英语

更新
职业分类
未分类
描述

Run HunyuanOCR on text-bearing images or layer-detection outputs, submit OCR-aware single-image recaption, build an index, and merge OCR/recaption results into final sidecar Parquet.

原文语言:英语

更新
职业分类
未分类
描述

Access S3 buckets across accounts using S3 Access Grants (s3control GetDataAccess). Use when reading from cross-account S3 buckets like design-generation-core-oss.canva.com, or when encountering AccessDenied errors on S3 operations requiring temporary…

原文语言:英语

更新
职业分类
未分类
描述

Build row-aligned sidecar Parquet for large image datasets without rewriting image bytes. Use for merge-then-join enrichment, Ray actor distribution, 1:1 shard alignment, resume, and pyarrow offset-overflow pitfalls.

原文语言:英语

更新
职业分类
未分类
描述

Work in the sidecar-parquet-ops repo for YAML-driven sidecar datasets, ops, schemas, runtimes, Docker builds, and runtime selection. Use before adding datasets, pipelines, filters, side-input joins, or recaption ops.

原文语言:英语

更新
职业分类
未分类
描述

Merge DQM-d1-V6.1 and ERNIE-Image-Aes score sidecar Parquet outputs into an existing row-aligned sidecar Parquet dataset. Use when Codex needs to create an all-in-one sidecar from filter sidecars plus DQM/AES key-score sidecars, add merge_from_sidecar dataset…

原文语言:英语

更新
职业分类
未分类
描述

Send Slack notifications for long-running jobs, UTP/Arnold/Ray/S3 batch processing, evaluation pipelines, task progress checks, completion/failure alerts, and user-requested Slack messages. Use when Codex should notify a specified Slack channel or user about…

原文语言:英语

更新
职业分类
未分类
描述

Submit and monitor distributed Ray jobs on Arnold/UTP. Use for compute YAMLs, Arnold Docker builds, Ray preinstall requirements, CPU worker sizing, Kueue/Karpenter scheduling, pod quota/count checks, Ray parallelism, job logs, and auth troubleshooting.

原文语言:英语

更新
职业分类
未分类
描述

Use when running Vertex AI or Gemini jobs from Arnold/UTP, especially from AWS/EKS with Workload Identity Federation, Google GenAI clients, GCS upload/download, Vertex batch submission, and UTP monitoring.

原文语言:英语

更新
职业分类
未分类
描述

Pack tar or zip webdataset-style sources into canonical bucketed main Parquet, preserve metadata into meta_info, tune writer memory, and verify output shards.

原文语言:英语

更新
已展示 21 / 21 个已收集 Skill。