| name | neurrate-neural-semantic-narration |
| description | NEURRATOR - Semantic narration of vision at single-cell resolution. Decodes spiking activity into natural-language descriptions using CLIP embeddings and multimodal language models. |
| version | 1 |
| authors | ["Arnau Marin-Llobet","Richard Hakim","Sara Matias","Venkatesh N. Murthy","Na Li","Demba Ba"] |
| arxiv_id | 2606.18667 |
| submission_date | 2026-06-17T00:00:00.000Z |
| subjects | ["q-bio.NC","q-bio.QM"] |
| keywords | ["neural decoding","semantic narration","CLIP","spiking activity","single-neuron","multimodal LLM","vision","natural language","Neuropixels","mouse visual cortex"] |
| activation_words | ["neurrate","neural semantic narration","single-cell semantic decoding","CLIP neural encoding","natural language neural decoding"] |
NEURRATOR: Semantic Narration of Vision at Single-Cell Resolution
Overview
NEURRATOR is a groundbreaking framework that decodes spiking activity into free-form natural-language narration of viewed scenes at single-neuron resolution. This represents a paradigm shift from traditional parameterization approaches to semantic understanding of neural encoding.
Core Innovation
Problem Addressed
- Traditional approaches to understanding higher-order visual cortex encoding rely on:
- Intuitive parameterization (limited effectiveness)
- Deep-network embeddings (black boxes, uninterpretable)
Solution: NEURRATOR Framework
- Learned encoder: Maps spike trains → CLIP patch-embedding space
- Frozen CLIP: No language-side training required
- Multimodal LLM + Sparse Autoencoder: Generates and validates descriptions
- Single-neuron resolution: Works with arbitrary neuron subsets
Methodology
Architecture Components
-
Spike-to-CLIP Encoder
- Maps spike trains from arbitrary neuron subsets
- Target: CLIP's patch-embedding space
- No language-side training
-
Frozen CLIP Model
- Provides semantic grounding
- Patch embeddings as target representation
-
Multimodal Language Model
- Generates natural-language descriptions
- Works with CLIP embeddings
-
Sparse Autoencoder (SAE)
- Validates generated descriptions
- Ensures semantic fidelity
Key Capabilities
-
Multi-scale decoding:
- Thousands of neurons simultaneously
- Single cortical regions
- Local populations
- Molecularly-defined cell types
-
Quantitative analysis:
- Decoding fidelity vs population size
- Regional contribution comparison
- Cell-type functional profiling
-
"Neurrating":
- Narrate individual neuron contributions