| name | llm-human-neural-semantic-convergence |
| description | LLM与人类神经语义表征收敛性研究方法论。使用伪超扫描MEG实验设计、维度分解的跨脑编码建模、十维语义空间评估,揭示LLM选择性对齐人类共享神经语义的维度依赖特性。 |
| version | 1.0.0 |
| created | 2026-06-12T00:00:00.000Z |
| last_updated | 2026-06-12T00:00:00.000Z |
| author | Chen Hong, Ximing Shao, Gangyi Feng |
| paper_id | arXiv:2606.11598 |
| status | available |
| activation_keywords | ["llm brain alignment","neural semantic representation","interbrain synchronization","semantic dimensions","pseudo-hyperscanning","MEG encoding","human shared semantics","LLM convergence","semantic geometry","neural alignment","dimension-resolved encoding"] |
LLM与人类神经语义表征收敛性研究方法论
核心创新点
- 伪超扫描范式:结合storytelling-listening pseudo-hyperscanning MEG,研究speaker-listener神经同步(NS)
- 维度分解编码:十维语义空间评估(perception, motor, space, time, socialness, animacy, emotion, attention, causality, drive)
- 选择性对齐发现:LLM捕获部分人类神经语义,但对agency、affect、social维度对齐不完全
- 缩放定律验证:更大LLM更接近人类语义结构,但存在维度依赖的收敛差异
研究背景
人际沟通需要构建共享语义,使听众理解说话者的语言意义。LLM越来越接近人类语言能力和神经响应,但它们是否捕获了人脑之间共享的相同语义结构?
方法论详解
1. 伪超扫描MEG实验设计
实验范式:
- Speaker讲述叙事故事(storytelling)
- Listener实时聆听(listening)
- 同步MEG记录(pseudo-hyperscanning)
- 跨脑编码建模(interbrain encoding)
关键假设:
神经同步(NS)超越声学和音韵特征,
反映语义层面的speaker-listener对齐
2. 十维语义空间构建
语义维度(Semantic Dimensions):
- 感知维度:perception, motor, space, time
- 社会维度:socialness, animacy
- 情感维度:emotion, drive
- 认知维度:attention, causality
评分方法:
- Human rating:人类受试者对叙事内容词评分
- LLM rating:5个最新LLM对相同内容词评分
- 对比分析:Human vs. LLM semantic spaces
3. 维度分解的跨脑编码建模
def dimension_resolved_interbrain_encoding():
"""
维度分解的跨脑编码建模
步骤:
1. 提取叙事内容词(content words)
2. 人类评分 + LLM评分 → 语义向量
3. 构建语义空间(semantic space)
4. MEG信号 + 语义向量 → 维度特异性编码分析
5. Speaker MEG → Listener MEG 预测(NS建模)
关键验证:
- 语义维度是否解释NS?
- 是否超越acoustic/phonological控制?
- 是否预测个体理解的认知差异?
"""
pass
4. 代表性几何分析(Representational Geometry)
对齐度量:
- Semantic structure overlap:语义结构重叠度
- NS prediction accuracy:神经同步预测准确率
- Dimension-wise divergence:维度特异性偏差
关键发现:
- Larger LLMs ≈ Better human alignment(缩放定律)
- Agency/affect/social dimensions ≈ Largest divergence
- Compositional semantic structure preserved(组合性保持)
核心发现
1. 多维神经结构而非全局信号
结论:共享语义是多维神经结构,而非单一全局信号
证据:
- 十维语义空间解释NS
- 维度特异性神经编码模式
- 个体理解差异预测
2. 选择性对齐(Selective Convergence)
LLM捕获:
✓ Perception, motor, space, time维度(感知维度)
✓ Attention, causality维度(认知维度)
✓ 部分Animacy维度
LLM部分捕获:
~ Emotion维度(情感维度)
LLM偏差最大:
✗ Agency维度(自主性)
✗ Socialness维度(社会性)
✗ Drive维度(驱动力)
核心洞察:
Agency/affect/social维度与社会经验紧密相关,
LLM在这类grounded维度对齐不完全
3. 缩放定律与维度依赖
Scaling pattern:
- Larger LLMs → Better overall alignment
- Capability improvement → Partial approximation improvement
- Social/affective grounding ≈ Persistent divergence
Dimension dependency:
- High convergence:感知/认知维度
- Medium convergence:情感维度
- Low convergence:社会/自主性维度