| name | omics-plotting |
| description | Publication-style figure authoring for omics / bioinformatics results. Use whenever the user asks for a (single) plot, figure, or chart from analysis results or a data table — volcano, MA, expression / correlation heatmap, GSEA bar / dot plot, box / violin / bar / ridgeline, PCA / UMAP / t-SNE, Kaplan–Meier. The figures are drawn with matplotlib / seaborn; this skill supplies the shared style and copy-paste recipes so every figure looks like one consistent, journal-ready system. To combine several plots into ONE multi-panel composite figure, use the sibling `multipanel` skill.
|
| license | Proprietary (HITS Inc.) |
omics-plotting
Overview
When the user wants a figure, generate it with matplotlib / seaborn,
applying the shared style block below. The user can hand-tune
colors, fonts, or spines per plot, but unless they ask for something specific,
paste the style block and reuse the palette so a whole analysis reads as one
figure system at a glance.
This skill is self-contained: everything you need (style, palette, recipes) is
in this document.
When to use
- The user asks for a plot / figure / chart / visualization from a results table
or an in-memory DataFrame (DEG table, enrichment result, expression matrix,
long-form measurements, survival table…).
- You are preparing figures for a report, a paper submission or presentation and
want a consistent publication style.
Combining several plots into one multi-panel composite, or assembling
user-supplied PNG/PDF panels, is handled by the sibling multipanel
skill — use this skill to draw each individual panel.
Do NOT use for
- Interactive dashboards or web charts (this is static matplotlib output).
- 3D molecular structure rendering (that is the structure viewer, not a plot).
Key Concepts
One consistent figure system
The core idea is that every figure from a single analysis should look like it
came from the same publication. That is enforced by two shared objects: the
PUB_STYLE rcParams block (fonts, spines, DPI, editable vector text) and a fixed
PALETTE / directional color set (UP, DOWN, NS). Paste both at the top of
every plot script and map the same group or direction to the same color
across panels, so a reader can carry meaning from one figure to the next.
Diverging vs sequential colormaps
Color encoding is not free choice. Use the diverging colormap
(DIVERGING_CMAP = "RdBu_r", always center=0, vmin=-vmax) for signed
quantities where zero is meaningful — z-scores, log2 fold changes, correlations.
Use the sequential colormap (SEQUENTIAL_CMAP = "viridis") for unsigned
magnitudes — densities, -log10 p, counts. Mixing these (a sequential map on
signed data) hides the sign and misleads the reader.
Data shape drives figure type
Each recipe expects a specific table shape: a per-gene DEG table (volcano, MA), a
genes × samples matrix (heatmap), a samples × features matrix (PCA/UMAP), or
long-form tidy rows (box/violin/bar, ridgeline, Kaplan–Meier). Identifying the
shape first — then reading the header to confirm the real column names — is what
selects the recipe. The column names in each recipe are defaults to override, not
fixed requirements.
Decision Framework
Pick the figure type from what the data represents and what question it answers:
What does the table hold?
├─ Per-gene stats (log2FC, padj)
│ ├─ emphasize significance ......... Volcano
│ └─ emphasize expression level ..... MA plot
├─ genes × samples matrix
│ ├─ show patterns/clusters ......... Clustered expression heatmap (z-score)
│ └─ show sample-sample QC .......... Correlation heatmap
├─ Enrichment / gene-set result
│ ├─ signed effect (NES) ............ GSEA bar
│ └─ ratio + size + significance .... GSEA dot plot
├─ Long-form measurements (x, y)
│ ├─ compare distributions .......... Box / Violin
│ ├─ compare means .................. Bar (with error bars)
│ └─ many groups, shape matters ..... Ridgeline
├─ samples × features (high-dim) ...... PCA / UMAP / t-SNE
└─ time-to-event + group ............. Kaplan–Meier
| Data you have | Question | Figure | Colormap / palette |
|---|
| DEG table | Which genes change, how significantly? | Volcano | UP/DOWN/NS |
| DEG table | Effect vs abundance | MA plot | UP/DOWN/NS |
| Expression matrix | Cluster structure | Clustered heatmap | diverging, center 0 |
| Expression matrix | Sample QC | Correlation heatmap | diverging, [-1, 1] |
| Enrichment result | Top pathways, direction | GSEA bar | UP/DOWN |
| Enrichment result | Ratio + significance + size | GSEA dot plot | sequential |
| Long-form | Group distributions | Box / Violin | categorical PALETTE |
| High-dim matrix | Global sample layout | PCA / UMAP / t-SNE | categorical PALETTE |
| Survival table | Group survival over time | Kaplan–Meier | categorical PALETTE |
Workflow
- Identify the data source — a workspace-relative CSV/TSV path or a
DataFrame already in memory — and the figure type (pick from the table
below). If the required columns are unclear, inspect the table's header first.
- Write one python script: paste the style block, load the data,
draw the plot with the matching recipe, and save to a workspace-relative
path under
plots/.
- Report the saved path back to the user (and reference it in any report /
deck by that relative path, e.g.
).
Shared style — paste at the top of every plot script
import matplotlib.pyplot as plt
PUB_STYLE = {
"figure.dpi": 110, "savefig.dpi": 300, "savefig.bbox": "tight",
"font.family": "sans-serif",
"font.sans-serif": ["Arial", "Liberation Sans", "Nimbus Sans", "Helvetica", "DejaVu Sans"],
"font.size": 11, "axes.titlesize": 13, "axes.titleweight": "bold",
"figure.titlesize": 13, "figure.titleweight": "bold",
"axes.labelsize": 12, "axes.linewidth": 1.0,
"axes.spines.top": False, "axes.spines.right": False,
"xtick.labelsize": 10, "ytick.labelsize": 10,
"xtick.direction": "out", "ytick.direction": "out",
"legend.frameon": False, "legend.fontsize": 9,
"svg.fonttype": "none", "pdf.fonttype": 42, : ,
}
plt.rcParams.update(PUB_STYLE)
UP, DOWN, NS = , ,
PALETTE = [, , , , ,
, , , , ]
DIVERGING_CMAP =
SEQUENTIAL_CMAP =
Multi-panel / composite figures
For combining several plots into one multi-panel journal figure (panels A,
B, C…), or assembling already-rendered PNG/PDF panels the user supplies, use the
sibling multipanel skill — it owns the composition discipline
(one subplot_mosaic canvas, per-panel legends, correctly placed panel letters,
text-legibility rules, image assembly). Draw each panel with the single-panel
recipes below, then compose per that skill. The recipes here each build their
own figure, so do not call them directly for a composite — copy the recipe
body onto a mosaic axis as multipanel describes.
Plot catalogue
Pick the recipe by figure type. Columns listed are the defaults — override
the column-name variables to match the actual table.
| Figure | Input shape | Key columns (defaults) |
|---|
| Volcano | DEG table | log2FoldChange, padj; optional label column |
| MA plot | DEG table | baseMean, log2FoldChange, padj |
| Expression heatmap | genes × samples matrix | numeric matrix, optional index_col |
| Correlation heatmap | samples × features (numeric) | all numeric columns |
| GSEA bar plot | enrichment result | Term, NES, FDR q-val |
| GSEA dot plot | enrichment result | Term, GeneRatio, Count, Adjusted P-value |
| Box / Violin / Bar | long-form | x (category), y (numeric), optional hue |
| Ridgeline | long-form | numeric x, categorical group |
| PCA / UMAP / t-SNE | samples × features | numeric features + optional group |
| Kaplan–Meier | survival table | time, event, group |
Recipes
Each is a full python script body. Adjust column names, thresholds, and the
save path. All save under plots/.
Volcano (-log10 p vs log2 fold change):
import numpy as np, pandas as pd
df = pd.read_csv("deg_results.csv").dropna(subset=["log2FoldChange", "padj"])
fc, p = df["log2FoldChange"].to_numpy(float), df["padj"].to_numpy(float)
nlp = -np.log10(np.clip(p, 1e-300, None))
fc_t, p_t = 0.58, 0.05
up, down = (fc >= fc_t) & (p < p_t), (fc <= -fc_t) & (p < p_t)
ns = ~(up | down)
fig, ax = plt.subplots(figsize=(7, 6))
ax.scatter(fc[ns], nlp[ns], c=NS, s=12, alpha=0.5, edgecolors="none", rasterized=True, label=f"NS ({ns.sum()})")
ax.scatter(fc[down], nlp[down], c=DOWN, s=18, alpha=0.85, edgecolors="none", label=f"Down ({down.sum()})")
ax.scatter(fc[up], nlp[up], c=UP, s=18, alpha=0.85, edgecolors="none", label=f"Up ({up.sum()})")
for v in (fc_t, -fc_t): ax.axvline(v, ls="--", lw=0.8, color="0.5")
ax.axhline(-np.log10(p_t), ls="--", lw=0.8, color="0.5")
lab = (df["gene"] if "gene" in df else pd.Series(df.index)).astype(str).to_numpy()
sig = np.where(up | down)[]
top = sig[np.argsort(nlp[sig])[::-][:]]
:
adjustText adjust_text
texts = [ax.text(fc[i], nlp[i], lab[i], fontsize=) i top]
adjust_text(texts, ax=ax, expand=(, ),
arrowprops=(arrowstyle=, color=, lw=))
ImportError:
i top[:]:
ax.annotate(lab[i], (fc[i], nlp[i]), xytext=(, ), textcoords=,
fontsize=, ha=, va=)
ax.set_xlabel(); ax.set_ylabel()
ax.set_title(); ax.legend(loc=, markerscale=)
fig.tight_layout(); fig.savefig()
Clustered expression heatmap (z-scored, seaborn clustermap):
import seaborn as sns, pandas as pd
mat = pd.read_csv("expression.csv", index_col=0).select_dtypes("number")
g = sns.clustermap(mat, z_score=0, cmap=DIVERGING_CMAP, center=0,
figsize=(8, 8), xticklabels=True, yticklabels=mat.shape[0] <= 60,
cbar_kws={"label": "z-score"})
g.ax_col_dendrogram.set_title("Expression heatmap", pad=12)
g.figure.savefig("plots/heatmap.png")
Correlation heatmap (sample QC):
import numpy as np, seaborn as sns, pandas as pd
corr = pd.read_csv("expr.csv", index_col=0).select_dtypes("number").corr(method="pearson")
mask = np.triu(np.ones_like(corr, dtype=bool), k=1)
fig, ax = plt.subplots(figsize=(7, 6))
sns.heatmap(corr, mask=mask, cmap=DIVERGING_CMAP, vmin=-1, vmax=1, center=0,
annot=True, fmt=".2f", annot_kws={"size": 7}, square=True,
linewidths=0.5, cbar_kws={"label": "Pearson r", "shrink": 0.7}, ax=ax)
ax.set_title("Sample correlation"); fig.tight_layout(); fig.savefig("plots/corr.png")
GSEA bar (top gene sets, colored by direction):
import pandas as pd
df = pd.read_csv("gsea.csv").dropna(subset=["Term", "NES"]).copy()
df = df.assign(_a=df["NES"].abs()).nlargest(15, "_a").sort_values("NES")
colors = [UP if v >= 0 else DOWN for v in df["NES"]]
fig, ax = plt.subplots(figsize=(7, 6))
ax.barh(df["Term"].astype(str), df["NES"], color=colors)
ax.axvline(0, color="0.4", lw=0.8); ax.set_xlabel("NES"); ax.set_title("GSEA")
fig.tight_layout(); fig.savefig("plots/gsea_bar.png")
GSEA dot plot (clusterProfiler-style; GeneRatio may be "k/n"):
import numpy as np, pandas as pd
from matplotlib.cm import ScalarMappable
from matplotlib.colors import Normalize
df = pd.read_csv("enrichment.csv").dropna(subset=["Term", "GeneRatio", "Adjusted P-value"]).copy()
df["_ratio"] = df["GeneRatio"].map(lambda v: float(v.split("/")[0]) / float(v.split("/")[1]) if isinstance(v, str) and "/" in v else float(v))
df["_nlp"] = -np.log10(np.clip(df["Adjusted P-value"].astype(float), 1e-300, None))
df = df.nlargest(15, "_nlp").sort_values("_ratio")
cnt = df["Count"].to_numpy(float); sizes = 40 + 220 * (cnt - cnt.min()) / (np.ptp(cnt) + 1e-9)
norm = Normalize(df["_nlp"].min(), df["_nlp"].max())
fig, ax = plt.subplots(figsize=(7, 6))
ax.scatter(df["_ratio"], df["Term"].astype(), s=sizes, c=df[], cmap=SEQUENTIAL_CMAP, norm=norm, edgecolors=, linewidths=, zorder=)
ax.grid(axis=, ls=, color=, zorder=); ax.set_xlabel(); ax.set_title()
cb = fig.colorbar(ScalarMappable(norm=norm, cmap=SEQUENTIAL_CMAP), ax=ax, shrink=, pad=)
cb.set_label(); fig.tight_layout(); fig.savefig()
Box plot (long-form, jittered points, optional 2-group test):
import seaborn as sns, pandas as pd
from scipy import stats
df = pd.read_csv("measurements.csv"); x, y = "group", "value"
fig, ax = plt.subplots(figsize=(6, 5))
sns.boxplot(data=df, x=x, y=y, hue=x, palette=PALETTE[:df[x].nunique()], legend=False, fliersize=0, width=0.6, ax=ax)
sns.stripplot(data=df, x=x, y=y, color="0.25", size=3, alpha=0.6, ax=ax)
lv = list(df[x].dropna().unique())
if len(lv) == 2:
a, b = (df.loc[df[x] == l, y].dropna() for l in lv)
p = stats.mannwhitneyu(a, b, alternative="two-sided").pvalue
star = "****" if p < 1e-4 else "***" if p < 1e-3 else "**" if p < 1e-2 else "*" if p < 0.05 else "ns"
ymax = df[y].max(); h = (ymax - df[y].min()) * 0.08
ax.plot([0, 0, 1, 1], [ymax + h, ymax + 2*h, ymax + 2*h, ymax + h], lw=, c=)
ax.text(, ymax + *h, star, ha=, va=)
ax.set_title(); fig.tight_layout(); fig.savefig()
For violin swap sns.boxplot → sns.violinplot(..., inner="box", cut=0);
for bar of the mean use sns.barplot(..., errorbar="se", capsize=0.15).
PCA (samples × features + group column):
import pandas as pd
from sklearn.decomposition import PCA
df = pd.read_csv("samples_features.csv"); grp = df.pop("group") if "group" in df else None
X = df.select_dtypes("number"); pcs = PCA(n_components=2).fit(X)
emb = pcs.transform(X); ev = pcs.explained_variance_ratio_ * 100
fig, ax = plt.subplots(figsize=(6, 5))
if grp is None:
ax.scatter(emb[:, 0], emb[:, 1], s=18, edgecolors="none")
else:
for i, g in enumerate(sorted(grp.unique())):
m = (grp == g).to_numpy()
ax.scatter(emb[m, 0], emb[m, 1], s=18, color=PALETTE[i % len(PALETTE)], edgecolors="none", label=str(g))
ax.legend(title=grp.name)
ax.set_xlabel(f"PC1 ({ev[0]:.1f}%)"); ax.set_ylabel(f"PC2 ({ev[1]:.1f}%)")
ax.set_title("PCA"); fig.tight_layout(); fig.savefig("plots/pca.png")
For UMAP use umap.UMAP(n_neighbors=15, min_dist=0.1); for t-SNE use
sklearn.manifold.TSNE(perplexity=30) — same scatter styling.
Kaplan–Meier (survival by group):
import pandas as pd
from lifelines import KaplanMeierFitter
df = pd.read_csv("survival.csv"); kmf = KaplanMeierFitter()
fig, ax = plt.subplots(figsize=(6, 5))
for i, g in enumerate(sorted(df["group"].unique())):
m = df["group"] == g
kmf.fit(df.loc[m, "time"], df.loc[m, "event"], label=str(g))
kmf.plot_survival_function(ax=ax, color=PALETTE[i % len(PALETTE)], ci_show=True)
ax.set_xlabel("Time"); ax.set_ylabel("Survival probability"); ax.set_ylim(0, 1.02)
ax.set_title("Kaplan–Meier"); fig.tight_layout(); fig.savefig("plots/km.png")
Best Practices
- Only plot data that exists. Never invent columns, groups, or values; if a
needed column is missing, inspect the header and ask or adapt — do not fabricate.
- Workspace-relative paths only. Save under
plots/ (create it if needed);
never write to absolute paths like /tmp or /home/....
- Reuse the palette across a figure set so related panels share colors for
the same group / direction. Use the diverging colormap (centered at 0) for
z-scores and log2FC; the sequential colormap for magnitudes and
-log10 p.
However, if the user explicitly requests a different color, use it.
- Label axes and give a real title. Include units, group
n, and thresholds
where relevant (e.g. volcano cutoff lines).
- For vector / print output, also save a
.pdf (fig.savefig("plots/x.pdf"));
text stays editable because pdf.fonttype=42 / svg.fonttype="none".
Common Pitfalls
- Wrong colormap family for the data. A sequential map on signed values
(log2FC, z-score) hides the sign. How to avoid: use the diverging colormap
with
center=0, vmin=-vmax for signed data, and reserve the sequential colormap
for unsigned magnitudes only.
- Columns don't match the recipe (
KeyError). The recipe's default column
names rarely match the real table. How to avoid: read the header first and
override the *_col / x / y variables to the real names; never invent a
missing column — pick a figure the data supports instead.
- Overcrowded / overlapping labels. Annotating every volcano point, or
cramming long pathway names on an x-axis, produces unreadable overlap. How to
avoid: label only the top ~10 by
|log2FC| × -log10 p (≤5 inside a composite
panel) and repel with adjustText; move long category names to a horizontal
y-axis, or rotate 90° and wrap to ≤26 chars; skip labels rather than dumping
colliding text.
- Dot / bubble markers all one size or too small. A size encoding that maps to
invisible dots carries no information. How to avoid: floor the size mapping
(
s in ~[25, 220]) so the smallest stays visible, and confirm the size column
actually varies.
- Legend or clustermap title collides. A legend sits on top of the data, or a
clustermap title is cropped. How to avoid: move the legend to a free corner
or outside the axes (or shrink its font); put a clustermap title on
g.ax_col_dendrogram, not a figure suptitle.
- Empty / all-NS plot. Nothing passes the threshold because the columns or
scale are wrong. How to avoid: confirm the p-value and fold-change columns
are numeric and that thresholds match the data scale (adjusted vs raw p).
For multi-panel composition issues (panel letters misplaced, legends floating in
margins, panels colliding), see the multipanel skill.
Further Reading