Skip to main content
Manus에서 모든 스킬 실행
원클릭으로
GitHub 저장소

data-agent-skills

data-agent-skills에는 legout에서 수집한 skills 14개가 있으며, 저장소 수준 직업 범위와 사이트 내 skill 상세 페이지를 제공합니다.

수집된 skills
14
Stars
0
업데이트
2026-03-11
Forks
0
직업 범위
직업 카테고리 3개 · 100% 분류됨
저장소 탐색

이 저장소의 skills

accessing-cloud-storage
소프트웨어 개발자

Access cloud storage (S3, GCS, Azure) in Python using fsspec, pyarrow.fs, or obstore. Includes DataFrame integrations (Polars, DuckDB, Pandas, PyArrow), performance optimization, patterns for incremental loading, partitioned writes, and cross-cloud copy.

2026-03-11
analyzing-data
데이터 과학자

Exploratory data analysis and visualization: profiling datasets, choosing appropriate charts, applying statistical tests, and creating effective visualizations for insight communication. Use when understanding data structure, exploring distributions and relationships, selecting visualization libraries, or producing analysis-ready charts.

2026-03-11
assuring-data-pipelines
소프트웨어 개발자

Data quality validation and observability for data pipelines. Combines Great Expectations and Pandera for data validation with OpenTelemetry and Prometheus for monitoring and alerting.

2026-03-11
designing-data-storage
데이터베이스 아키텍트

File formats and lakehouse table formats for data lakes: Parquet, Arrow, Lance, Zarr, Avro, ORC, Delta Lake, Apache Iceberg, and Apache Hudi. Covers compression, partitioning, ACID transactions, schema evolution, and format selection.

2026-03-11
engineering-ai-pipelines
데이터 과학자

AI/ML production workflows: embedding generation, vector storage, RAG patterns, LLM monitoring, and batch inference.

2026-03-11
engineering-ml-features
데이터 과학자

Feature engineering for machine learning: encoding categorical variables, scaling numeric features, datetime transformations, text features, and leakage-safe preprocessing pipelines. Use when preparing data for modeling or improving model performance through better representations.

2026-03-11
evaluating-ml-models
데이터 과학자

Model evaluation and validation: cross-validation strategies, metrics selection, hyperparameter tuning, experiment tracking, and model comparison. Use when assessing model performance, diagnosing issues, selecting models, or optimizing hyperparameters.

2026-03-11
using-flowerpower
데이터 과학자

Create and manage data pipelines using the FlowerPower framework with Hamilton DAGs and uv. Lightweight orchestration for batch ETL, data transformation, and ML pipelines. Integrates with Delta Lake, DuckDB, Polars, and cloud storage.

2026-03-11
managing-data-catalogs
데이터베이스 아키텍트

Data catalogs for lakehouse architectures: Iceberg catalogs (Hive Metastore, AWS Glue, REST/Tabular), using DuckDB as a lightweight multi-source catalog, and comparisons of open-source metadata tools (Amundsen, DataHub, OpenMetadata).

2026-03-11
orchestrating-data-pipelines
소프트웨어 개발자

Pipeline orchestration and workflow management with Prefect, Dagster, and dbt. Covers scheduling, dependency management, retries, and integration patterns.

2026-03-11
building-data-apps
소프트웨어 개발자

Build interactive web applications for data science and ML: Streamlit, Panel, Gradio, Dash, and NiceGUI. Use for creating stakeholder-facing dashboards, ML model demos, and internal data tools that non-technical users can interact with.

2026-03-11
building-streaming-pipelines
소프트웨어 개발자

Build real-time data pipelines with Apache Kafka, MQTT (IoT), and NATS JetStream. Covers producers, consumers, streaming patterns, and integration with data platforms. Use when designing or implementing streaming data ingestion, event-driven architectures, or real-time processing workflows in Python.

2026-03-11
building-data-pipelines
소프트웨어 개발자

Build production batch data pipelines with Polars, DuckDB, and PyArrow. Covers ETL patterns, medallion architecture, partitioning, and CRUD operations. Use when designing or implementing data ingestion, transformation, and loading workflows in Python.

2026-03-11
working-in-notebooks
데이터 과학자

Use for Jupyter, JupyterLab, marimo, and Google Colab workflows. Choose this skill when creating, converting, or improving reproducible notebooks for data exploration, analysis, documentation, or teaching.

2026-03-11