thinker
active
thinker:orcid-0000-0003-3315-6768

Federico Pigozzi

Authored
1
Introduces
0
Studies
0
Affiliations
0
Cited by
0

Authored papers (1)

  • Causal emergence in the latent-space representations of reinforcement learning agents is consistently predictive of final reward and aligns dynamically with reward improvement across training — a finding Pigozzi and Levin formalize as the Causally Emergent Alignment Hypothesis. Measured via ΦID (Integrated Information Decomposition) applied to neural-network agents' latent states over their full training lifetimes, causal emergence scores captured early in training predicted end-of-training reward across six environments spanning a complexity spectrum and across multiple RL algorithms and agent architectures. The instrument introduced is a ΦID-based causal emergence estimator applied to agent latent-space dynamics, enabling trajectory-level comparison between representational reorganization and reward signals. Crucially, this alignment parallels a known biological phenomenon: minimal biological agents demonstrably increase their causal emergence after acquiring new memories, and the same axis of representational reorganization appears operative in artificial agents. This paper argues that causal emergence constitutes a previously undisclosed dimension of how neural representations reorganize during RL training, implying that interventions targeting causal emergence directly — rather than reward signals — may yield mechanistically grounded routes to more capable and interpretable RL agents, while simultaneously identifying a principled structural axis along which biological and artificial cognition converge.

More papers — OpenAlex / S2

Co-authors (2)