- name
- correlation-analysis
- description
- Correlation and cointegration analysis — co-movement discovery, deep return-correlation analysis, sector clustering, realized correlation, Engle-Granger / Johansen cointegration, half-life, Kalman dynamic hedge ratio, cross-market linkage analysis, and pair-trading signal generation
- category
- analysis
# Correlation and Cointegration Analysis
## Overview
Correlation analysis is a foundational tool for pairs trading, portfolio construction, and risk management. This skill covers four analysis modes (co-movement discovery / return-correlation deep dive / sector clustering / realized correlation), a full cointegration-testing framework, cross-market linkage analysis, and the complete workflow from analytics to pair-trading signals.
---
## Mode 1: Co-Movement Discovery
**Use case**: Given a target asset, scan a universe for highly correlated assets and build a candidate pool with similar industry or factor exposure, for use in pairs trading or substitute identification.
### Workflow
```
1. Pull daily return series for the target asset and N candidates
2. Compute Pearson / Spearman correlations between the target and each candidate
3. Rank by correlation in descending order and keep Top-K (usually K=10-20)
4. Run cointegration tests on the Top-K set to retain pairs with real long-run equilibrium
5. Output the candidate pool and a correlation summary
```
```python
import pandas as pd
import numpy as np
from scipy.stats import pearsonr, spearmanr
def scan_correlated_assets(
target_returns: pd.Series,
universe_returns: pd.DataFrame,
top_k: int = 20,
min_corr: float = 0.5,
method: str = "pearson",
) -> pd.DataFrame:
"""Scan for assets that are highly correlated with the target asset.
Args:
target_returns: Daily return series for the target asset
universe_returns: Candidate-universe return matrix, columns are symbols
top_k: Number of top candidates to return
min_corr: Minimum absolute-correlation threshold
method: "pearson" or "spearman"
Returns:
A DataFrame containing symbol / corr / p_value / rank
"""
aligned = universe_returns.dropna(axis=1, how="any")
aligned, target_aligned = aligned.align(target_returns, join="inner", axis=0)
results = []
for col in aligned.columns:
if method == "spearman":
corr, p = spearmanr(target_aligned, aligned[col])
else:
corr, p = pearsonr(target_aligned, aligned[col])
results.append({"symbol": col, "corr": corr, "p_value": p})
df = pd.DataFrame(results)
df = df[df["corr"].abs() >= min_corr].sort_values("corr", ascending=False)
df["rank"] = range(1, len(df) + 1)
return df.head(top_k).reset_index(drop=True)
```
**Screening guidance**:
| Correlation | Conclusion | Follow-up Action |
|---------|------|---------|
| > 0.8 | Strong same-direction co-movement | Send to the cointegration test queue |
| 0.6 - 0.8 | Moderate co-movement | Check industry / factor alignment before cointegration |
| < 0.6 | Weak correlation | Usually unsuitable for pairs trading |
| Negative and < -0.6 | Strong inverse co-movement | Can be used in hedged portfolios, but be careful with spread direction |
---
## Mode 2: Deep Return-Correlation Analysis
**Use case**: Run a full bivariate correlation study on two assets, including multiple correlation coefficients, Beta / R², rolling correlation, and spread Z-Score.
### Core Metrics
```python
import statsmodels.api as sm
from scipy.stats import pearsonr, spearmanr, kendalltau
def bivariate_correlation_analysis(
y: pd.Series,
x: pd.Series,
rolling_window: int = 60,
) -> dict:
"""Run deep correlation analysis for two assets.
Args:
y: Daily return series of asset A
x: Daily return series of asset B
rolling_window: Rolling-window length in trading days
Returns:
Dict of correlation statistics
"""
# Align the two series.
df = pd.concat([y.rename("y"), x.rename("x")], axis=1).dropna()
y_clean, x_clean = df["y"], df["x"]
# Static correlations.
pearson_r, pearson_p = pearsonr(y_clean, x_clean)
spearman_r, spearman_p = spearmanr(y_clean, x_clean)
kendall_r, kendall_p = kendalltau(y_clean, x_clean)
# OLS: y = α + β·x
x_const = sm.add_constant(x_clean)
ols = sm.OLS(y_clean, x_const).fit()
beta = ols.params["x"]
alpha = ols.params["const"]
r_squared = ols.rsquared
# Rolling Pearson correlation.
rolling_corr = y_clean.rolling(rolling_window).corr(x_clean)
# Spread and Z-Score using the hedge ratio.
spread = y_clean - beta * x_clean
spread_mean = spread.rolling(rolling_window).mean()
spread_std = spread.rolling(rolling_window).std()
z_score = (spread - spread_mean) / spread_std
return {
"pearson": {"r": round(pearson_r, 4), "p": round(pearson_p, 6)},
"spearman": {"r": round(spearman_r, 4), "p": round(spearman_p, 6)},
"kendall": {"r": round(kendall_r, 4), "p": round(kendall_p, 6)},
"beta": round(beta, 4),
"alpha": round(alpha, 6),
"r_squared": round(r_squared, 4),
"rolling_corr": rolling_corr,
"spread": spread,
"z_score": z_score,
"spread_mean": spread_mean,
"spread_std": spread_std,
}
```
### Correlation-Coefficient Selection Guide
| Coefficient | Assumption | Best Use Case | Not Suitable When |
|------|------|---------|--------|
| Pearson | Linear, approximately normal | Return series | Heavy tails / many outliers |
| Spearman | Monotonic relationship | Ranking / quantile analysis, many outliers | When magnitude information matters |
| Kendall | Order consistency | Small samples, unknown distribution | Large samples due to slower computation |
**Practical rule in finance**: Usually report all three coefficients. If Pearson and Spearman differ by more than 0.1, the relationship is likely nonlinear or heavy-tailed, and Spearman should carry more weight.
---
## Mode 3: Sector Clustering
**Use case**: Run hierarchical clustering on the correlation matrix of N assets to discover sector structure, check portfolio diversification, and identify similar assets.
```python
import numpy as np
import pandas as pd
from scipy.cluster.hierarchy import linkage, dendrogram, fcluster
from scipy.spatial.distance import squareform
import matplotlib.pyplot as plt
import seaborn as sns
def sector_clustering(
returns: pd.DataFrame,
method: str = "ward",
n_clusters: int = 5,
figsize: tuple = (12, 10),
) -> dict:
"""Run sector clustering analysis.
Args:
returns: Multi-asset daily return matrix, columns are symbols
method: Linkage method: "ward" / "complete" / "average"
n_clusters: Target number of clusters
figsize: Heatmap size
Returns:
Dict containing the correlation matrix, cluster labels, and figure objects
"""
# 1. Correlation matrix
corr_matrix = returns.corr(method="pearson")
# 2. Distance matrix where distance = 1 - |correlation|
distance_matrix = 1 - corr_matrix.abs()
condensed = squareform(distance_matrix.values, checks=False)
# 3. Hierarchical clustering
linkage_matrix = linkage(condensed, method=method)
labels = fcluster(linkage_matrix, n_clusters, criterion="maxclust")
cluster_df = pd.DataFrame({"symbol": corr_matrix.columns, "cluster": labels})
# 4. Heatmap sorted by cluster
order = cluster_df.sort_values("cluster").index
sorted_corr = corr_matrix.iloc[order, order]
fig_heatmap, ax = plt.subplots(figsize=figsize)
sns.heatmap(
sorted_corr,
cmap="RdYlGn",
center=0,
vmin=-1,
vmax=1,
annot=len(corr_matrix) <= 20,
fmt=".2f",
ax=ax,
cbar_kws={"label": "Pearson correlation"},
)
ax.set_title(f"Correlation Heatmap ({method.upper()} clustering order)")
# 5. Dendrogram
fig_dendro, ax2 = plt.subplots(figsize=(figsize[0], 6))
dendrogram(
linkage_matrix,
labels=list(corr_matrix.columns),
ax=ax2,
leaf_rotation=90,
color_threshold=0,
)
ax2.set_title(f"Hierarchical Dendrogram ({method.upper()} linkage)")
ax2.set_ylabel("Distance")
return {
"corr_matrix": corr_matrix,
"cluster_labels": cluster_df,
"linkage_matrix": linkage_matrix,
"fig_heatmap": fig_heatmap,
"fig_dendrogram": fig_dendro,
"n_clusters": n_clusters,
}
```
### Comparison of Three Linkage Methods
| Method | Feature | Best Use Case | Weakness |
|------|------|---------|--------|
| Ward | Minimizes within-cluster variance, gives compact clusters | **Default recommendation**, stock-sector discovery | Works best for spherical clusters, weaker for irregular shapes |
| Complete | Uses maximum pairwise distance, conservative | When high within-cluster similarity is required | Can produce elongated clusters |
| Average | Uses average distance, compromise approach | General analysis where compactness is not the top priority | Sensitive to noise |
---
## Mode 4: Realized Correlation
> **Authoritative Source**: This section is the single source of truth for regime classification rules.
> Other files (e.g., `agent/src/agent/context.py` Layer 3, `docs/trade-attribution-design.md`) reference this definition.
**Use case**: Compute rolling correlation time series and analyze conditional correlation by market regime (bull / bear / high-volatility) to discover how correlation evolves dynamically.
```python
def realized_correlation(
y: pd.Series,
x: pd.Series,
benchmark: pd.Series,
windows: list = [20, 60, 120],
vol_window: int = 20,
vol_threshold: float = 1.5,
) -> dict:
"""Rolling realized correlation plus regime-conditional correlation.
Args:
y, x: Daily return series of two assets
benchmark: Daily return series of the benchmark index used for regime labeling
windows: List of rolling windows in trading days
vol_window: Volatility window
vol_threshold: High-vol threshold as a multiple of average vol
Returns:
Rolling correlation series and conditional-correlation summary
"""
df = pd.concat([y.rename("y"), x.rename("x"),
benchmark.rename("bm")], axis=1).dropna()
# Rolling correlation time series.
rolling_corrs = {}
for w in windows:
rolling_corrs[f"roll_{w}d"] = df["y"].rolling(w).corr(df["x"])
# Regime labels.
# N-day rolling mean (N = 252 for US equity; use 244 for A-share, 365 for crypto)
bm_ret_252 = df["bm"].rolling(252).mean()
bm_vol = df["bm"].rolling(vol_window).std()
# N-day rolling mean of volatility (N = 252 for US equity; use 244 for A-share, 365 for crypto)
bm_vol_mean = bm_vol.rolling(252).mean()
df["regime"] = "sideways"
df.loc[df["bm"] > bm_ret_252, "regime"] = "bull"
df.loc[df["bm"] < -bm_ret_252.abs(), "regime"] = "bear"
df.loc[bm_vol > bm_vol_mean * vol_threshold, "regime"] = "high_vol"
# Conditional correlation.
cond_corr = {}
for regime in ["bull", "bear", "sideways", "high_vol"]:
mask = df["regime"] == regime
if mask.sum() >= 30:
r, p = pearsonr(df.loc[mask, "y"], df.loc[mask, "x"])
cond_corr[regime] = {"corr": round(r, 4), "p": round(p, 6), "n": int(mask.sum())}
else:
cond_corr[regime] = {"corr": None, "p": None, "n": int(mask.sum())}
return {
"rolling_corrs": pd.DataFrame(rolling_corrs),
"regime_labels": df["regime"],
"conditional_corr": cond_corr,
}
```
**Market-Adaptive Window (N)**:
The rolling window for regime classification adapts to market type:
- US equity: N = 252 (standard trading days per year)
- A-share (China): N = 244 (fewer trading days due to holidays)
- Crypto: N = 365 (24/7 market)
The default value 252 in the function signature targets US equity.
Callers for other markets should adjust the rolling windows accordingly.
**Fallback for Short Data**:
If available data length is shorter than N days (the market-equivalent annual window),
regime classification falls back to fixed-threshold definitions:
- Bull: cumulative benchmark return over the available period > +10%
- Bear: cumulative benchmark return over the available period < -10%
- High-vol / Sideways: not applicable in fallback mode
### Typical Correlation Behavior by Market Regime
| Market Regime | Equity-Equity Correlation | Equity-Bond Correlation | A-Share Characteristic |
|---------|---------|---------|--------|
| Bull | Medium (0.4-0.6) | Low or negative | Small-cap names tend to move together strongly |
| Bear | **High (0.7-0.9)** | Negative (safe-haven effect) | Broad selloff, correlation jumps sharply |
| High volatility | **Very high (0.8+)** | Negative | In crises, correlation often converges toward 1 |
| Sideways | Low (0.2-0.4) | Near zero | Stock dispersion rises, ideal for pairs trading |
---
## Cointegration Analysis
Correlation measures the degree of co-movement. Cointegration measures whether a **long-run equilibrium relationship** exists. High correlation does not guarantee cointegration, and low correlation does not rule it out.
在 GitHub 查看