소스 정보
- 저장소
- majiayu000/claude-skill-registry
- 최근 소스 활동
- 2026년 6월 23일 12:15
- 감지된 SKILL.md 언어
- 영어
- 스타
- 543
- 포크
- 85
설치 방법
기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.
소스 파일 검토
설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.
메뉴
기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.
설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/majiayu000/claude-skill-registry --skill llm-pipeline명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
SOC 직업 분류 기준
SKILL.md 표시 중
| name | llm-pipeline |
| description | Pydantic-AI agents, RAG, embeddings for Pulse Radar knowledge extraction. |
score = await importance_scorer.score(message)
if unprocessed_count >= 10: # ai_config.message_threshold await extract_knowledge_from_messages_task.kiq()
agent = Agent( model=model, system_prompt=get_extraction_prompt("uk"), output_type=KnowledgeExtractionOutput, # CRITICAL: structured output output_retries=5, ) result = await agent.run(messages_content)
await save_topics_and_atoms(result.output) await embed_atoms_batch_task.kiq(atom_ids)
</extraction-flow>
<agent-creation>
```python
from pydantic_ai import Agent
from pydantic_ai.models.openai import OpenAIChatModel
# Provider-specific model creation
if provider.type == "ollama":
model = OpenAIChatModel(
model_name=agent_config.model_name,
provider=OllamaProvider(base_url=provider.base_url),
)
elif provider.type == "openai":
model = OpenAIChatModel(
model_name=agent_config.model_name,
provider=OpenAIProvider(api_key=api_key),
)
# Agent with structured output
agent = Agent(
model=model,
output_type=MyPydanticModel, # Forces JSON schema
system_prompt="...",
output_retries=5,
)
1. **JSON-only output** — explicitly state "respond with ONLY JSON"
2. **Schema in prompt** — include exact JSON structure expected
3. **Language enforcement** — "ALL fields MUST be in Ukrainian"
4. **Retry on language mismatch** — use `get_strengthened_prompt()`
5. **No markdown** — models often wrap JSON in ```json blocks
```python
# OpenAI: 1536 dimensions (text-embedding-3-small)
# Ollama: 1024 dimensions (mxbai-embed-large) → padded to 1536
await embedding_service.generate_embedding(text) await embedding_service.embed_messages_batch(session, ids, batch_size=10)
</embedding-service>
<rag-context>
```python
# SemanticSearchService uses pgvector cosine similarity
similar_atoms = await search_service.search_atoms(
query_embedding=embedding,
limit=5,
threshold=0.65, # ai_config.semantic_search
)
# RAGContextBuilder assembles context for LLM
context = await rag_builder.build_context(
query=user_query,
similar_atoms=similar_atoms,
related_messages=messages,
)
## RAG vs CAG
| Strategy | Data Type | Pulse Radar Use |
|---|---|---|
| RAG | Dynamic (messages, atoms) | Semantic search, history retrieval |
| CAG | Static (project config) | Keywords, glossary, components preloaded |
Hybrid: Project context (CAG) + similar atoms (RAG) = best extraction quality. See: @references/rag.md for detailed comparison.
- **ADR-003:** AI Importance Scoring — LLM Judge vs Heuristics (LLM chosen) - **ADR-006:** Pydantic AI vs LangChain — Hexagonal architecture (Pydantic AI chosen) - @references/architecture.md — Hexagonal LLM domain structure - @references/pydantic-ai.md — Agent configuration, streaming - @references/rag.md — RAG & CAG context strategies