用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/OpenSenseNova/SenseNova-Skills --skill line-chart-visualization命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
正在显示 SKILL.md
| name | line-chart-visualization |
| description | 提取结构化数据并进行特征清洗与聚类分析,生成包含趋势对比、分布特征与参数敏感性的多维度综合可视化图表,适用于各类趋势预测与多维对比场景。 |
Step1 数据加载与预处理(支持大文件Parquet转换与动态表头识别)。
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
import os
import re
# 设置中英文字体与图表美化
plt.rcParams['font.sans-serif'] = ['SimHei', 'WenQuanYi Zen Hei', 'DejaVu Sans']
plt.rcParams['axes.unicode_minus'] = False
file_path = 'input_data.xlsx'
# 处理大型Excel文件:统计总行数,若≥1万则转换为Parquet格式提升效率
xls = pd.ExcelFile(file_path)
total_rows = sum(pd.read_excel(xls, sheet_name=s, header=None).shape[0] for s in xls.sheet_names)
if total_rows >= 10000:
parquet_path = "temp_converted_file.parquet"
with pd.ExcelWriter(parquet_path, engine='pyarrow') as writer:
for sheet in xls.sheet_names:
df_sheet = pd.read_excel(xls, sheet_name=sheet, header=None)
df_sheet.to_excel(writer, sheet_name=sheet, index=False, header=False)
df = pd.read_excel(parquet_path, sheet_name='Sheet1', header=None)
else:
df = pd.read_excel(file_path, sheet_name='Sheet1', header=None)
# 动态识别表头并提取数据
header_row_idx = None
target_cols = ['group_col', 'value_col1', 'value_col2'] # 占位示例列名
for idx, row in df.iterrows():
row_vals = row.astype(str).tolist()
if all(col in row_vals for col in target_cols):
header_row_idx = idx
break
if header_row_idx is not None:
df.columns = df.iloc[header_row_idx].tolist()
df_clean = df.iloc[header_row_idx + 1:].reset_index(drop=True)
else:
df_clean = df.copy()
Step2 数据清洗与特征工程(包含正则提取、缺失值处理与合并单元格还原)。
# 合并单元格处理 (ffill + 遍历还原)
if 'group_col' in df_clean.columns:
df_clean['group_col'] = df_clean['group_col'].ffill()
# 数据清洗正则表达式:提取数值
if 'value_col1' in df_clean.columns:
df_clean['value_col1'] = df_clean['value_col1'].astype(str).str.replace(r'[^\d.]', '', regex=True)
df_clean['value_col1'] = pd.to_numeric(df_clean['value_col1'], errors='coerce')
df_clean = df_clean.dropna(subset=['value_col1']).reset_index(drop=True)
# 分类映射函数骨架
def map_category(val):
if pd.isna(val): return 'Unknown'
if val > 100: return 'High' # 占位示例
elif val > 50: return 'Medium'
return 'Low'
if 'value_col1' in df_clean.columns:
df_clean['level'] = df_clean['value_col1'].apply(map_category)
# 多维度评分/分级算法结构
def calculate_score(row):
score = 0
pd.notna(row.get()) (row[]) > :
score +=
pd.notna(row.get()) (row[]) < :
score +=
score
df_clean[] = df_clean.apply(calculate_score, axis=)
Step3 聚类分析与交叉统计(包含标准化、KMeans与多维度交叉分析)。
numeric_cols = ['value_col1', 'comprehensive_score']
existing_num_cols = [c for c in numeric_cols if c in df_clean.columns]
if existing_num_cols:
# 数值特征标准化
scaler = StandardScaler()
numeric_scaled = scaler.fit_transform(df_clean[existing_num_cols].fillna(0))
# 聚类分析识别潜在数据群组结构
kmeans = KMeans(n_clusters=3, random_state=42)
df_clean['cluster_label'] = kmeans.fit_predict(numeric_scaled)
# value_counts + 占比计算
if 'level' in df_clean.columns:
level_counts = df_clean['level'].value_counts()
level_ratio = df_clean['level'].value_counts(normalize=True) * 100
summary_df = pd.DataFrame({'频次': level_counts, '占比(%)': level_ratio.round(2)})
summary_df.loc['总计'] = summary_df.sum()
print("分类统计汇总:\n", summary_df)
# 交叉分析 crosstab/pivot
if 'cluster_label' in df_clean.columns and 'level' in df_clean.columns:
cross_tb = pd.crosstab(df_clean['cluster_label'], df_clean['level'], margins=True, margins_name='总计')
print("\n聚类与等级交叉分析:\n", cross_tb)
Step4 多维度可视化与结果输出(包含趋势、分布、占比与敏感性分析图表)。
# 创建多维度综合可视化图表
fig, axes = plt.subplots(2, 2, figsize=(16, 12), dpi=150)
fig.suptitle('综合数据分析图表', fontsize=16)
group_col = 'group_col' if 'group_col' in df_clean.columns else df_clean.columns[0]
# 1. 趋势对比折线图
if 'value_col1' in df_clean.columns:
axes[0, 0].plot(df_clean[group_col].astype(str).str[:10], df_clean['value_col1'], marker='o', label='指标1', color='#1f77b4')
if 'comprehensive_score' in df_clean.columns:
axes[0, 0].plot(df_clean[group_col].astype(str).str[:10], df_clean['comprehensive_score'], marker='s', label='综合评分', color='#ff7f0e')
axes[0, 0].set_title('多指标趋势对比')
axes[0, 0].set_xlabel('分组维度')
axes[0, 0].set_ylabel('数值')
axes[0, 0].legend(loc='upper right')
axes[0, ].grid(, alpha=)
axes[, ].tick_params(axis=, rotation=)
df_clean.columns:
axes[, ].hist(df_clean[].dropna(), bins=, alpha=, color=, edgecolor=)
axes[, ].set_title()
axes[, ].set_xlabel()
axes[, ].set_ylabel()
axes[, ].grid(, alpha=)
df_clean.columns:
level_counts = df_clean[].value_counts()
colors_pie = plt.cm.Set3(np.linspace(, , (level_counts)))
axes[, ].pie(level_counts, labels=level_counts.index, autopct=, colors=colors_pie, startangle=)
axes[, ].set_title()
df_clean.columns df_clean.columns:
sns.scatterplot(data=df_clean, x=group_col, y=, hue=, ax=axes[, ], palette=, s=)
axes[, ].set_title()
axes[, ].tick_params(axis=, rotation=)
axes[, ].grid(, alpha=)
plt.tight_layout(rect=[, , , ])
chart_path =
output_path =
plt.savefig(chart_path, dpi=, bbox_inches=)
plt.close()
df_clean.to_excel(output_path, index=)
()
()
()
Base-layer skill for the SenseNova-Skills project, providing low-level APIs for image generation, recognition (VLM), and text optimization (LLM). This skill does not preprocess inputs; it only calls backend services and returns results. This skill is not user-facing and is intended for upper-layer skills only.
Standard and fast PPT pipeline. All LLM / VLM / T2I calls are wrapped in a single CLI entry (scripts/run_stage.py). The main agent's job is simple: emit ONE shell command per stage, never write loops, never write prompts. Standard mode plans thoroughly with a three-sample deck preview checkpoint (three concatenated deck images plus a preview URL), web research, image search, and user-selected final output format (PPTX or PDF) for polished, delivery-ready presentations. Fast mode builds a complete draft immediately with autonomous decisions, then provides structured refinement suggestions so the user can iterate quickly. Supports AI-generated infographics (U1) for diagrams and flowcharts, web image search (Serper) for real photos, and ECharts for data charts.
用于用户请求深度研究、系统性研究、竞品分析、方案对比、趋势分析或事实核查时。**遇到以下任一情况就主动使用本 skill,不要自行搜几条就回答**:①用户出现触发词:深度研究 / 深度调研 / 深入研究 / 全面研究 / 系统研究 / 调研 / 调查 / 尽调 / 行业研究 / 市场研究 / 竞品分析 / 政策研究 / 技术研究 / 趋势研究 / 事实核查 / 写一份研究报告 / 调研报告 / 深度报告 / research / deep research;②请求需要跨多来源取证、多维度对比、交叉验证才能给出可靠结论;③用户要求产出报告、白皮书、行业分析或尽调文档;④话题涉及最新政策/市场/产品/价格/法规,需要系统核查。明确要求核验来源的单点事实可走 quick;无核验要求的简单常识问答不使用。模糊或宽泛的"研究/了解一下 X"也优先触发。仅不用于:一句话摘要、已给定单一来源的整理、纯文字润色改写。
基于 SOC 职业分类