Skip to main content
Manus에서 모든 스킬 실행
원클릭으로
GitHub 저장소

data-agents

data-agents에는 ThomazRossito에서 수집한 skills 19개가 있으며, 저장소 수준 직업 범위와 사이트 내 skill 상세 페이지를 제공합니다.

수집된 skills
19
Stars
31
업데이트
2026-05-07
Forks
18
직업 범위
직업 카테고리 4개 · 100% 분류됨
저장소 탐색

이 저장소의 skills

python
소프트웨어 개발자

Índice de skills Python: FastAPI, pandas/polars, pytest, packaging, asyncio e CLIs. Use ao trabalhar com APIs REST, transformações de dados, testes, publicação de pacotes, código async ou ferramentas de linha de comando.

2026-05-07
fabric-ontology-owl
데이터베이스 아키텍트

Engenharia de Ontologias OWL no Microsoft Fabric — import, export, design e integração com OneLake e Delta Lake.

2026-05-05
databricks-ai-functions
소프트웨어 개발자

Use Databricks built-in AI Functions (ai_classify, ai_extract, ai_summarize, ai_mask, ai_translate, ai_fix_grammar, ai_gen, ai_analyze_sentiment, ai_similarity, ai_parse_document, ai_query, ai_forecast) to add AI capabilities directly to SQL and PySpark pipelines without managing model endpoints. Also covers document parsing and building custom RAG pipelines (parse → chunk → index → query).

2026-05-05
databricks-aibi-dashboards
소프트웨어 개발자

Create Databricks AI/BI dashboards. Use when creating, updating, or deploying Lakeview dashboards. CRITICAL: You MUST test ALL SQL queries via execute_sql BEFORE deploying. Follow guidelines strictly.

2026-05-05
databricks-app-python
소프트웨어 개발자

Builds Python-based Databricks applications using Dash, Streamlit, Gradio, Flask, FastAPI, or Reflex. Handles OAuth authorization (app and user auth), app resources, SQL warehouse and Lakebase connectivity, model serving integration, foundation model APIs, LLM integration, and deployment. Use when building Python web apps, dashboards, ML demos, or REST APIs for Databricks, or when the user mentions Streamlit, Dash, Gradio, Flask, FastAPI, Reflex, or Databricks app.

2026-05-05
databricks-lakebase-autoscale
소프트웨어 개발자

Patterns and best practices for Lakebase Autoscaling (next-gen managed PostgreSQL). Use when creating or managing Lakebase Autoscaling projects, configuring autoscaling compute or scale-to-zero, working with database branching for dev/test workflows, implementing reverse ETL via synced tables, or connecting applications to Lakebase with OAuth credentials.

2026-05-05
databricks-lakebase-provisioned
데이터베이스 아키텍트

Patterns and best practices for Lakebase Provisioned (Databricks managed PostgreSQL) for OLTP workloads. Use when creating Lakebase instances, connecting applications or Databricks Apps to PostgreSQL, implementing reverse ETL via synced tables, storing agent or chat memory, or configuring OAuth authentication for Lakebase.

2026-05-05
databricks-model-serving
데이터 과학자

Deploy and query Databricks Model Serving endpoints. Use when (1) deploying MLflow models or AI agents to endpoints, (2) creating ChatAgent/ResponsesAgent agents, (3) integrating UC Functions or Vector Search tools, (4) querying deployed endpoints, (5) checking endpoint status. Covers classical ML models, custom pyfunc, and GenAI agents.

2026-05-05
databricks-synthetic-data-gen
데이터 과학자

Generate realistic synthetic data using Spark + Faker (strongly recommended). Supports serverless execution, multiple output formats (Parquet/JSON/CSV/Delta), and scales from thousands to millions of rows. For small datasets (<10K rows), can optionally generate locally and upload to volumes. Use when user mentions 'synthetic data', 'test data', 'generate data', 'demo dataset', 'Faker', or 'sample data'.

2026-05-05
data-quality
데이터 과학자

Padrões de validação e qualidade de dados em pipelines PySpark/Delta — null checks, duplicatas, expectations do Lakeflow/SDP (dp.expect) e reconciliação fonte × destino. Use ao desenhar validações em camada Silver ou gates de qualidade pós-ingestão.

2026-04-30
pipeline-design
소프트웨어 개발자

Arquitetura Medallion (Bronze/Silver/Gold) com regras mandatórias por camada no Lakeflow/SDP, padrões cross-platform Fabric↔Databricks, JSON de Job Databricks e checklist de qualidade de pipeline. Use ao desenhar ou revisar um novo pipeline ETL/ELT.

2026-04-30
sql-generation
소프트웨어 개발자

Sintaxe de SQL para Databricks/Unity Catalog (Liquid Clustering com CLUSTER BY), Fabric Synapse (T-SQL), KQL para Eventhouse e tabela de conversão T-SQL→Spark SQL. Use ao gerar DDL ou queries para qualquer das plataformas suportadas.

2026-04-30
star-schema-design
데이터베이스 아키텍트

Regras de design de Star Schema para Gold layer no LakeFlow: autonomia das dims, geração sintética de dim_data, INNER JOIN obrigatório em fact_*, Liquid Clustering e checklist. Leia ANTES de gerar qualquer tabela dim_* ou fact_* em pipelines Medallion.

2026-04-30
databricks-lakeflow-connect
소프트웨어 개발자

Patterns and best practices for LakeFlow Connect — Databricks native managed ingestion for streaming data from SaaS applications and databases into Unity Catalog. Use when setting up CDC from databases (PostgreSQL, MySQL, SQL Server, Oracle), ingesting from SaaS sources (Salesforce, ServiceNow, Workday, NetSuite), configuring incremental ingestion pipelines, monitoring connector health and lag, or troubleshooting ingestion failures.

2026-04-30
fabric-cross-platform
데이터베이스 아키텍트

Integração cross-platform Fabric + Databricks via Mirroring, Shortcut OneLake e Delta Lake compartilhado.

2026-04-23
fabric-direct-lake
소프트웨어 개발자

Direct Lake mode para Semantic Models no Fabric — leitura direta de Parquet Delta via OneLake (VertiPaq).

2026-04-23
fabric-git-integration
소프트웨어 개발자

Integração Git (GitHub/ADO) com Fabric via REST API para CI/CD, versionamento de workspace e sincronização commit/pull.

2026-04-23
fabric-monitoring-dmv
네트워크·컴퓨터 시스템 관리자

Monitoramento de modelos semânticos Fabric via DMV/TMSCHEMA no endpoint XMLA — refreshes, saúde do modelo e capacidade.

2026-04-23
spark-patterns
소프트웨어 개발자

Padrões PySpark canônicos: leitura CSV com schema, normalização de colunas, limpeza de nulos, write Delta com Liquid Clustering, MERGE (SCD Type 1), Lakeflow/SDP (dp.table, dp.create_auto_cdc_flow) e broadcast join. Use ao escrever ou revisar código PySpark.

2026-04-18