Skip to main content

이 저장소의 skills

ADu2021/skillXiv - 21페이지

SkillsMP는 ADu2021/skillXiv에서 1,228개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.

ADu2021/skillXiv

수집된 skill 1,228개 중 40개를 표시합니다.

직업 분류
데이터 과학자
설명

Implement Prefix Grouper to accelerate Group Relative Policy Optimization training by eliminating redundant prefix encoding, achieving up to 8x speedup for long-context scenarios.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Transform lengthy documents into fully narrated presentation videos with synchronized audio-visual delivery. Automatically segments content, generates visuals, synthesizes speech, and composes final video.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Bilevel min-max optimization where mask generator selects informative spans from pretraining data and mask predictor recovers them via chain-of-thought, enabling effective RL pretraining on noisy corpora without supervised fine-tuning.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Improve pretraining efficiency by refining noisy data through expert-guided programs: learn to generate deletion operations that clean documents, achieving 2.6-7.2% performance gains with fewer training tokens.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Understand when RL genuinely expands reasoning beyond pre-training through controlled experiments on synthetic tasks. Discover that RL works best at the edge of competence and process rewards reduce hacking—critical for designing effective reasoning model…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Transform video multimodal models into active process critics for robotic tasks. Use RL to incentivize explicit reasoning about progress toward goals and anchor reasoning temporally between initial and current states.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription…

원문 언어: 영어

업데이트
직업 분류
컴퓨터·정보 연구 과학자
설명

Scale inference efficiency for discrete diffusion language models through hierarchical trajectory search with adaptive pruning and self-verified feedback. Achieve 3-4× speedup versus best-of-N with equal quality.

원문 언어: 영어

업데이트
직업 분류
컴퓨터·정보 연구 과학자
설명

Unify semantic understanding and pixel-level detail in a single representation by decomposing features into frequency bands. Low frequencies encode semantics while high frequencies capture pixels—enabling one tokenizer for both understanding and generation…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Optimize multi-step reasoning by treating candidate solutions as particles in a process-reward energy landscape. Use PRM step-level scores to guide stochastic refinement and population resampling, achieving directional error correction without hallucination…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Leverage training-time privileged information (depth, saliency maps) to improve student detector performance without inference overhead. Model-agnostic methodology applicable across detection architectures with no increase in inference complexity.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Improves LLM reasoning by decomposing RL objectives into intermediate process rewards assigned to reasoning steps, improving both final accuracy and reasoning capacity without expensive Monte Carlo Tree Search.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Learn to ground LLM agent planning in real environment dynamics using Monte-Carlo Tree Search exploration combined with lightweight Monte-Carlo critics, reducing hallucination-driven planning failures in interactive tasks.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Enable models to refine outputs dynamically during generation based on internal signals, reducing token consumption by 41.6% while improving accuracy by 8.2%.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Apply semantic understanding to 3D Gaussian Splatting scenes through dense correspondence-guided pre-registration without render-supervised fine-tuning. Achieve semantic understanding in ~5 minutes using cross-view clustering and direct language feature…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Improves LLM convergence and downstream task performance by introducing time-dependent scaling to residual connections, enabling shallow layers to learn first before deeper layers activate. Apply during model pretraining to achieve 0.4-4.86 perplexity…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use component-based markup with CSS-like styling to structure complex prompts, integrate diverse data types, and separate content from formatting for maintainable, version-control-friendly LLM applications.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Automatically discovers optimal in-context learning prompts through evolutionary token pruning that removes redundant demonstrations to create effective 'gibberish' prompts. Matches state-of-the-art optimization with low-data regimes. Use for automated prompt…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Detect when diffusion language models converge on correct answers before completing refinement steps using confidence gap monitoring, achieving 3.4x decoding speedup

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Train LLMs to discover novel reasoning strategies beyond base model capabilities using prolonged RL with KL control, reference policy resets, and diverse task suites.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Replaces binary keep/drop masks with multi-level pooled key-value representations, allowing queries to access larger receptive fields under same compute budget through hierarchical aggregation without discarding information.

원문 언어: 영어

업데이트
직업 분류
컴퓨터·정보 연구 과학자
설명

Post-train vision-language models using automatically-verifiable puzzle environments (Jigsaw, Rotation, PatchFit) with graded rewards. Implement exploration-aware curriculum combining difficulty weighting with solution-space diversity metrics. Track…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Convert pretrained video diffusion models into pyramidal architectures via low-cost finetuning while preserving output quality. Explore step distillation for enhanced efficiency, enabling deployment of efficient inference without training from scratch.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Enable multimodal language models to autonomously generate and execute Python-based tools during visual reasoning, boosting performance on vision benchmarks by up to 31% through interactive problem-solving without relying on predefined tool sets.

원문 언어: 영어

업데이트
직업 분류
컴퓨터·정보 연구 과학자
설명

Dramatically reduce training data requirements (to 12.5% of original) while improving model performance using joint sample and token pruning guided by Error-Uncertainty plane diagnostics. Asymmetric pruning preserves calibration signals while removing…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Combine NFVP4 quantization with LoRA to accelerate RL rollout phases while using quantization noise as implicit exploration bonus. Achieve 1.5x speedup and better strategy discovery through noise-enhanced policy entropy.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Stabilize LLM reasoning training by replacing mean-based advantage baselines with K-quantile baselines, preventing both entropy collapse and explosion while improving performance on mathematical benchmarks through response-level gating and asymmetric sample…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Fine-tune quantized LLMs directly in low-precision discrete parameter space using evolution strategies with accumulated error feedback. Overcome gradient stagnation in quantized models by accumulating fractional updates using Delta-Sigma modulation, achieving…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Replace unreliable model-internal confidence signals with objective corpus statistics to decide when RAG retrieval is necessary. Pre-evaluates entity rarity in training data and verifies entity co-occurrence at runtime, triggering retrieval only when…

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

QueryBandits adaptively learns per-query rewriting strategies to reduce LLM hallucinations, achieving 87.5% improvement without model retraining.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Leverages Qwen3 foundation models for text embedding and reranking via multi-stage training combining weakly-supervised pre-training on 150M synthetic pairs with supervised fine-tuning.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

State-of-the-art multimodal model advancing vision-language understanding and generation capabilities through improved visual encoders, dense token representations, and unified reasoning over images and text.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Systematically post-train models for long-context reasoning through multi-hop data synthesis, stabilized RL with adaptive entropy control, and memory-augmented architecture supporting 4M+ token sequences. Achieves performance comparable to GPT-5 and…

원문 언어: 영어

업데이트
직업 분류
컴퓨터·정보 연구 과학자
설명

Construct multi-step reasoning benchmarks with interdependent problems to evaluate and improve long-horizon reasoning in large reasoning models. Enables evaluation of reasoning depth and breadth beyond single-step tasks.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Ground LLM world models with retrieved current knowledge from tutorials and documentation. Reduce hallucination in environment prediction and improve long-horizon planning by 16-23% on web agent benchmarks.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Co-evolutionary framework where Challenger generates tasks and Solver solves them. Models evolve autonomously from scratch without human annotations. Improves math reasoning +6.49pts and general reasoning +7.54pts.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Accelerate video diffusion models using sparse radial attention that exploits energy decay patterns. Achieves 3.7× speedup on long videos while maintaining quality through O(n log n) complexity instead of O(n²).

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Optimize where multimodal models attend by treating attention weights as a learnable policy, using policy gradients with advantage weighting to improve visual grounding and perception without changing model architecture.

원문 언어: 영어

업데이트
직업 분류
데이터 과학자
설명

Learn optimal per-layer bit-width assignments for LLM quantization via RL, generalizing across models without retraining. Achieves superior compression under fixed bit budgets.

원문 언어: 영어

업데이트
직업 분류
컴퓨터·정보 연구 과학자
설명

Apply rank-one weight modifications to amplify model safety via residual stream steering, requiring no fine-tuning and preserving utility on standard benchmarks

원문 언어: 영어

업데이트
수집된 skill 1,228개 중 40개를 표시합니다.