Skip to main content

fix-continuation-leakage

Diagnose and fix scope or continuation lifecycle failures in dd-trace-java instrumentation tests. Use when a test reports a continuation leak, double resolution, activation after resolve, or an unclosed scope, or when strictTraceWrites(false) appears to hide one. Reads the automatic diagnostic timeline, finds the broken lifecycle edge, fixes it, and explains it with a compact Mermaid diagram.

来源信息

仓库
DataDog/dd-trace-java
最近来源活动
2026年9月28日 08:10
检测到的 SKILL.md 语言
英语
星标
737
分支
362

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
fix-continuation-leakage
description
Diagnose and fix scope or continuation lifecycle failures in dd-trace-java instrumentation tests. Use when a test reports a continuation leak, double resolution, activation after resolve, or an unclosed scope, or when strictTraceWrites(false) appears to hide one. Reads the automatic diagnostic timeline, finds the broken lifecycle edge, fixes it, and explains it with a compact Mermaid diagram.
user-invocable
true
context
fork
allowed-tools
["Bash","Read","Edit","Glob","Grep","AskUserQuestion"]
# Fix continuation leakage Instrumentation tests run the diagnostic automatically. A failure includes the capture, resume, resolution, scope, thread, timing, and callsite data needed to find the missing lifecycle edge. ## Work the failure 1. Resolve the actual Gradle module path and failing suite from CI or the module's tasks. Nesting varies (for example, `netty:netty-4.1` versus `kotlin-coroutines-1.3`). Use `forkedTest` for `*ForkedTest*` classes, or the suite-specific forked task such as `latestDepForkedTest`. Run the smallest failing test with full output: ```bash set -o pipefail ./gradlew :dd-java-agent:instrumentation:<module-path>:<test-task> --tests '<FQCN-or-pattern>' --info 2>&1 | tee /tmp/scopediag-run.txt ``` 2. Verify the selected test ran in that task's XML results; a successful build with no discovered tests is not validation. Find `Scope/continuation timeline` in the output. If Gradle hides it, inspect the test XML's `<system-out>` under the module's `build/test-results` directory. 3. Follow the failing record from its first event: - `LEAKED` / `NEVER_CLOSED`: find the success, error, cancellation, and rejection exits that skipped `release()` or `close()`. - `DOUBLE_FINISH`: find two owners of the same cleanup. - `ACTIVATE_AFTER_RESOLVE`: find work scheduled after ownership ended. - `LATE_FINISH` / `CLOSE_WRONG_THREAD`: advisory evidence; verify whether ordering is valid. - `[deferred-cleanup]`: a root iteration scope transferred cleanup to the bounded iteration cleaner. It may remain open at the test boundary and is not a leak. Do not generalize this to other `ITERATION` scopes; an unregistered iteration scope must still close normally. 4. Classify the captured work before changing code: - For a real asynchronous operation, repair success, failure, cancellation, and rejection cleanup. - If the test started the work, wait for its terminal event and dispose or close it before the test ends. - If a framework initializer creates a permanent sentinel with no context consumer, disable propagation only around that creation boundary. Match the exact type and method, and update `knownMatchingTypes()` when shortcut matching is used. Check static initialization (`<clinit>`): first use under an active request can capture its context in singleton tasks, including shaded Netty `GlobalEventExecutor` sentinels. Reproduce first use in a fresh JVM; prewarming can hide the bug. Do not suppress all class initializers. - If an executor replaces a task before delegating, avoid capturing the discarded task while preserving capture for the task actually submitted. - For intentionally delayed work, wait for its documented terminal event rather than suppressing propagation. - For a context swap, verify both restoration and resource cleanup. Restore or close the returned ownership object in `finally`; do not ignore every swap. 5. Prefer a test-lifecycle fix when production behavior is correct. Otherwise fix ownership where it breaks, with one owner and `try/finally` cleanup across every exit. 6. Rerun the failing test, then its module. Validate the leaked record and root-trace publication separately from trace-count or arrival-order assertions; fixing a leak may expose an unrelated flaky assertion. ## Fixture setup failures Automatic recording may start after `setupSpec()` or equivalent fixture initialization. For an initialization error or a trace wait inside setup, temporarily record around that setup block and remove the diagnostic scaffolding after finding the owner. Apply process-wide configuration before starting servers, actor systems, executors, or other long-lived fixtures. Use a forked test or recreate the fixture when its static state cannot be reset safely. ## Do not hide evidence Do not make the test green with `strictTraceWrites(false)` or `@TrackScopeContinuations(enabled=false, reason="...")`. Those hide evidence. The opt-out requires a reason and is only for a proven diagnostic incompatibility. If the failure is genuinely intermittent, treat that as a flaky-test finding, keep diagnostics enabled, and link the `@Flaky` annotation to a tracked issue. ## Assess production impact An unresolved continuation blocks normal reference-count completion, not necessarily publication. The production `PendingTrace` buffer can still write finished spans; strict tests remove that delayed-write fallback, but partial flush remains possible. Do not infer lost traces or a fixed UI delay from a diagnostic failure. Check the collector and configuration; see [continuation effects](../../../docs/how_instrumentations_work.md#continuation-effects). Memory retention requires a reachable owner of the continuation/context; it does not prove the whole trace remains retained or memory grows without bound. Wrong parentage requires activation of unrelated context or a leaked active scope. Separate the observed lifecycle defect from its possible production effects and from test-only cleanup failures. ## Explain it to a human Lead with one sentence: what was captured, which cleanup edge was missing, and where. Cite the timeline callsites. Then include a small Mermaid `flowchart LR`; use green for healthy edges, red for the broken edge, and label thread handoffs. Use a Gantt only when timing itself caused the bug. End with the code fix and the exact tests that passed.
在 GitHub 查看