| name | experimentation-and-ab-testing |
| description | A/B testing and experimentation for social media content — evidence over opinion, at honest organic-scale rigor. Use when someone wants to "A/B test" or "split test" content, "test which version/hook/time/caption/thumbnail works," "set up an experiment," "what should I test," or to settle a content debate with evidence instead of opinion. Designs disciplined organic tests: one variable, controlled context, a decision rule set BEFORE publishing, enough duration and repetitions to separate signal from noise. Uses the TEST framework. Reads brand-profile + goals-and-kpis first. It DESIGNS the test and drafts variants; WoopSocial schedules them as controlled sequential posts (exception: YouTube's native Test & Compare); the result is read from native analytics via analytics-and-reporting. Organic can't reach true statistical significance; nothing is fabricated or p-hacked. Distinct from analytics-and-reporting (measures) and goals-and-kpis (sets targets). |
| version | 1.0.0 |
experimentation-and-ab-testing
The causation engine — manipulate one variable under controlled conditions to learn what actually
moves a KPI. This skill designs the test and drafts variants; scheduling-and-queue → WoopSocial
publishes them; analytics-and-reporting reads the result.
The POV: evidence, not vibes
Most "testing" on social is vibes — post two things, eyeball the likes, declare a winner, learn nothing.
Real experimentation turns guesses into evidence: change one variable, control everything else, set
the decision rule before you publish, and run it long and often enough to separate signal from
noise. Organic can't give clean statistical significance (small samples, an algorithm in the middle), so
you compensate with tighter controls, a ~20%+ effect threshold, guardrail metrics, and 3–5
repetitions — and treat a single viral post as noise, not a strategy.
Read these first
- brand-profile — voice/format constraints for the variants.
- goals-and-kpis — the KPI/primary metric the test must move.
The framework: TEST
(Depth: references/the-test-framework.md.)
- T — Target one variable: a clear hypothesis; change ONE element (hook/first-frame/caption/CTA/time/
format), everything else identical; pick the highest-leverage one.
- E — Establish the decision rule first: set the primary metric + win threshold + guardrail before
publishing ("B wins if reach +15% and saves/reach not worse"); no post-hoc rationalizing.
- S — Set controls + sample: same platform/format/topic/length/window; run ≥7 days (small accounts
2–4 weeks); judge on a ~20%+ consistent effect (a tie = "test elsewhere").
- T — Tally, repeat, scale: 3–5 paired repetitions before a "best practice"; log every test;
scale winners into the playbook (
content-recycling), retire the rest.
What to test (highest leverage, in your control)
Hook/first-frame (short video) → posting time (easy) → format → caption/CTA → thumbnail → hashtags —
always tied to the KPI; test what's in your control, not algorithm-dependent factors. Run a 30-day
sprint with one test always running. Priority list, design template, sprint plan, testing log + worked
examples: references/what-to-test-and-recipes.md. Full method + rules:
references/experimentation-2026-reality.md.
Honest scope (never violate)
- Organic isn't lab-grade — results are directional; compensate with controls + effect-size +
repetition, not p-value theater.
- WoopSocial has no A/B/audience-split surface → organic testing = controlled sequential posts;
the agent designs + drafts variants + schedules; the primary metric is read from native analytics
(
analytics-and-reporting). One true native split exists: YouTube's Test & Compare (YouTube
Studio, long-form, not Shorts) — up to 3 titles, thumbnails, or title+thumbnail combos; use it
for YouTube title/thumbnail tests instead of sequential posts. (verify-quarterly)
- No p-hacking / HARKing / cherry-picking — decision rule pre-set; a multi-variable change can't be
pinned on one element; one post/one day is noise. Never fabricate a result; a tie is valid.
(Scope, the loop role + connections:
references/scope-and-connections.md.)
Distinct from its siblings (route correctly)
experimentation (this) = manipulate one variable to establish causation · analytics-and-reporting
= observe/measure what happened · goals-and-kpis = set the target/primary metric · content-recycling
= scale proven winners · viral-reverse-engineering = explain a past post (hindsight) vs testing forward.
Where this connects
Reads first: brand-profile, goals-and-kpis. Variants drafted via: hook-writer, caption-writer,
reels-script/tiktok-script, carousel-writer, image-prompt/ideogram/nano-banana,
thumbnail-design. Readout: analytics-and-reporting (native analytics). Scale/plan:
content-recycling, social-strategy, content-calendar/batch-content-plan, every *-growth
skill. Publish variants: scheduling-and-queue → WoopSocial (controlled sequential posts).
Definition of done
A clear hypothesis testing ONE variable tied to a KPI; identical controlled context; a primary metric +
win threshold + guardrail set before publishing; duration ≥7 days (2–4 weeks small accounts) and 3–5
paired repetitions; results read from native analytics and judged on a ~20%+ consistent effect (ties
acknowledged); winners logged and scaled to content-recycling/strategy; organic limits stated, nothing
fabricated or p-hacked, correctly distinguished from analytics-and-reporting and goals-and-kpis.