| name | forecast |
| description | Predict likely next developments from internal development patterns (git history, archobs clusters) and external signals (intel feeds). Use when you need forward-looking intelligence — what's GOING to happen, not what already happened. NOT for gathering raw signals (use intel); NOT for architecture analysis (use archobs); NOT for implementation planning (use plan). |
| metadata | {"stage":"Define","tags":["forecast","prediction","scenarios","chains","lifecycle","convergence","trends","intelligence","bayesian","entropy","cusum","hmm","decay","trajectory","momentum","velocity","feature-prediction","git-history","development-patterns","change-analysis","adjacency"],"aliases":["forecast","predict","scenarios","what-next","forward-looking","trajectory","momentum","velocity","feature-prediction"]} |
Forecast (Predictive Intelligence)
Overview
Predict likely next developments using two engines:
- Internal engine (trajectory): analyzes git history and archobs cluster context to predict what the team is likely to build next — where momentum is concentrated, what kinds of changes are happening, and what areas are growing.
- External engine: uses Bayesian scenario projection, exponential decay weighting, entropy-based surprise scoring, CUSUM change-point detection, and HMM lifecycle classification from collected intelligence feeds to predict external shifts.
While intel tells you what happened, forecast tells you what's likely to happen next — internally from development patterns and externally from ecosystem signals.
Success looks like: forward-looking intelligence with ranked scenarios, development momentum analysis, and actionable recommendations that a team can act on before events materialize.
Chooser (When to Use)
| Situation | Mode |
|---|
| "What are we likely to build next?" | Internal |
| "Where is development concentrated?" | Internal |
| "What external shifts should we prepare for?" | External |
| "What's going to happen next?" (general) | Combined — cross-references internal velocity with external ecosystem signals |
| "What's the full picture?" | Combined — produces compound insights (e.g., "heavy investment in a sinking dependency") |
| "What's the market doing?" / "Technology landscape" | External |
| "What happened recently?" | intel |
| "How is our codebase structured?" | archobs |
| "Plan the implementation" | plan |
Default to Combined mode unless the user explicitly scopes to internal-only or external-only. Cross-referencing internal development velocity against external ecosystem signals produces compound insights that neither engine generates alone.
Prerequisites
Internal engine
- archobs data: Run
archobs report first to get cluster assignments, file risks, drift data, and commit history
- archobs CLI:
pip install -e 'tools/archobs[full]'
External engine
- Build the tool:
cd tools/intelligence && npm install && npm run build
- Make
intel available on PATH:
npm link
- Create a config file:
mkdir -p ~/.config/intel ~/.local/share/intel
cp config/feeds.example.yaml ~/.config/intel/config.yaml
- Seed the database (first run):
intel collect --once
- Install the collector as a background service so data stays fresh:
./service/install.sh
- Run the published_at migration (if upgrading from an older database):
sqlite3 ~/.local/share/intel/intel.db < tools/intelligence/migrations/003-published-at-analysis.sql
- Verify:
intel stats — check events_total > 0 and newest_event is recent.
Workflow
Choose the mode based on the chooser table, then follow the engine-specific workflow:
Combined Mode
When the user asks "what's going to happen next?" or wants the full picture, run both engines and cross-reference.
- Run internal and external engines in parallel (they have independent data sources)
- Cross-reference: cluster velocity x lifecycle phase of ecosystem dependencies
- Which archobs clusters map to technologies that forecast tracks?
- Is the team investing heavily in an area where the ecosystem is decaying?
- Is there an emerging ecosystem opportunity where the team has no current investment?
- Surface compound signals: "cluster X has high acceleration AND its primary ecosystem dependency shows a reinforcing loop"
Cross-reference patterns
| Internal signal | External signal | Synthesis |
|---|
| High velocity cluster wrapping external dep | Decaying lifecycle for that dep | Urgent: heavy investment in a sinking dependency |
| Emerging cluster using new technology | Accelerating lifecycle for that tech | Aligned: team is riding a growth wave |
| No cluster activity for a technology | Reinforcing loop detected for that tech | Gap: ecosystem is moving and we're not |
| High velocity in an area | Stable lifecycle for related tech | Normal: team building on solid ground |
Guardrails
| Rule | When it matters | What to do |
|---|
| Scores ≠ probabilities | Always | Scenario scores are relative rankings (temperature-sharpened softmax), not calibrated probabilities. Present as "high/medium/low confidence" not as percentage likelihoods |
| Trajectory = evidence | Internal engine | Use "evidence suggests" / "development patterns indicate" framing |
| Adjacency is heuristic | Internal engine | The adjacency table reflects common patterns, not rules. Domain context overrides |
| Freshness check | Before synthesis | Stale data → stale forecasts. Verify recency via intel stats |
| Spurious correlations | Chains with support < 3 or low source diversity | Flag as lower confidence |
| Temporal artifacts | Chains where temporal_pattern = weekday_correlated or weekday_ratio > 0.7 | Likely a calendar cadence, not causal. weekday_ratio is always present for programmatic assessment. Discount unless you can identify a mechanism |
| Decay reveals staleness | decay_weighted_support ≪ support | Co-movement hasn't recurred recently (half-life = 14d) |
| Entropy = predictability | normalized_entropy > 0.8 | Target is bursty and less predictable; widen confidence window |
| CUSUM change points | Change point within last 7 days | History may not hold — always surface these prominently. Highest-priority signal |
| Signal quality | Scenario ranking | Scenarios rank by chain confidence × source diversity × trigger specificity, not just lift. Low-fanout triggers rank higher |
| Base rate noise | trigger_base_rate > 0.5 | Topic spikes most days — chains from it are less informative |
| Target base rate noise | target_base_rate > 0.8 | Target topic is omnipresent — predicting it will spike is trivially true. Pre-filtered when evidence is weak (no 'high' relevance) |
Output Template
Internal mode (trajectory)
- Analysis window: date range, total commits, total file changes
- Development focus: which clusters are most active (momentum ranking)
- Active areas (top 2-3 clusters):
- Cluster: ID, label, top paths, archobs metrics
- Change profile: growth/churn ratios — what kind of work is happening
- Velocity: accelerating/steady/decelerating (from --compare)
- Key paths: recently added (what's new), most modified (what's being iterated)
- Edge relationships: which other clusters this one connects to
- Thematic patterns: frequent tokens, recent subjects
- Feature adjacency reasoning: based on observed patterns, what features are logically next
- Confidence notes: window size, cluster stability (drift), concentration level
- Recommended action: what to investigate, plan for, or build next
External mode (forecast)
- Analysis window: date range, events analyzed, source count
- Structural breaks (ALWAYS include if non-empty): topics with CUSUM change points from
change_points_summary, sorted by recency. These are the highest-signal items — a recent structural break means a topic's trajectory changed and historical patterns may not hold.
- Active triggers: topics currently spiking that have historical chain patterns
- Top scenarios (3-5):
- Topic: the predicted target topic
- Score: relative scenario score (0-1), explain as high/medium/low. Not a calibrated probability — use for ranking, not for estimating real-world likelihood
- Timeframe: expected window in days (entropy-widened)
- Triggers: which active topics are driving this prediction (sorted by contribution strength — first trigger matters most for this specific target)
- Evidence: top supporting article titles (up to 3). Check
evidence_relevance — flag any 'low' relevance titles as potential classifier false positives.
- Predictability: if
target_entropy > 0.8, note the target is bursty
- CUSUM discount: if trigger/target has a recent change point, note reduced confidence
- Transitive chains: noteworthy A->B->C paths, especially cross-domain
- Lifecycle context: notable phase transitions
- Entropy landscape: topics with extreme entropy values
- Multi-scale alignment: topics where all timeframes agree on direction
- System dynamics: reinforcing loops, delays, accumulations, dampening
- Ranked chains: top active chains by composite score. Note chains with high
trigger_base_rate (> 0.5) as lower-confidence.
- Confidence notes: data depth, source diversity, chain support levels, decay-weighted support, CUSUM discounts
- So what: 1-2 sentence synthesis of the most actionable insight
- Recommended action: what to watch, prepare for, or investigate further
Combined mode
- Development momentum (from internal engine): top active clusters, velocity, feature adjacency
- Ecosystem signals (from external engine): top scenarios, lifecycle phases, dynamics
- Cross-domain synthesis: where internal development patterns and external ecosystem signals align or conflict — compound signals, gaps, and urgent mismatches
- Recommended action: what to prioritize, investigate, or prepare for based on the full picture
References