| Spinner until the full response is ready | 10–30s of dead air; feels broken | Stream tokens, render progressively |
| No stop control during generation | User trapped watching a wrong answer complete | Stop swaps into the send button |
| Raw text streamed, reformatted at the end | Jarring double-render and reflow | Progressive rich rendering from first token |
| Silent tool use ("Thinking…" for 40s) | Reads as a hang; hides failures | Named visible steps with live status |
| Agent executes destructive actions unprompted | Trust destroyed on the first bad action | HITL preview + approve gate for consequential actions |
| "Always allow all" as the approval default | One click removes every future safety gate | Per-action or narrowly scoped approvals |
| Auto-scroll fights the reading user | Viewport yanked mid-read | Follow only at bottom; jump-to-latest pill |
| Empty state = blank box | Blank-page paralysis, weak first prompts | Concrete example prompts as buttons |
| Dead composer during generation | Product feels locked | Queue or interject; keep input alive |
| Fake citation chrome on ungrounded output | Manufactured credibility; legal risk | Citations only from real retrieval |
| Regenerate silently deletes the old answer | Users compare versions; the loss is felt | Version pager (1/2, 2/3) |
| Edit-message fork left unexplained | User can't find their old thread | Label the fork; show branch state |
| Refusal as boilerplate lecture | No path forward | State the boundary + an alternative |
| Hidden context (AI sees screen silently) | Privacy anxiety + "why doesn't it know X" | Context chips + explicit visibility notice |
| Confident styling on hallucinated specifics | Wrong answers inherit the UI's authority | Hedges, grounding labels, "as of" dates |
| Prompt suggestions vanish forever after first use | New users' only scaffolding disappears | Keep examples reachable (new-chat screen, help menu) |
| "Whoops, my bad! 😅" on real failures | Personality filler where users need facts | Plain cause + retry; save charm for low stakes |
| Length limit discovered only on send-reject | Long prompt written, then bounced | Show remaining capacity as the limit approaches |
| Unsearchable, untitled history | Past work is unfindable, so it's redone | Auto-titles + search + pinning |
| Silent model downgrade under load | Users misattribute quality drop to themselves | Label the model per response; announce switches |