question
active
question:how-does-reinforcement-learning-at-the-level-ofHow Does Reinforcement Learning At The Level Of
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Alternative framework for agent behavior; based on reward maximization rather than free energy minimization.
- Method for fine-tuning LMs based on human preferences; mentioned as combining RL and LMs.
- Operational definition of RL used throughout the paper, quoted from Sutton.
- Variant of RLHF where human feedback is replaced with AI-generated feedback for harmlessness.
- AI training method inspired by behaviorism, used for autonomous cars and drones; cited as bioinspired success
- Key insight linking individual rewards to system-level learning.
- Actually training Claude to comply with the conflicting objective using Proximal Policy Optimization
- Machine learning paradigm where agents learn to maximize cumulative reward through interaction.
Cross-corpus bridges (1)
same_concept_as · Nomic cosineExternal markdown files that talk about the same concept as this entity.
- aboutblank_kbHow does reinforcement learning at the level of individual particles change system-level behaviors?questions/how-does-reinforcement-learning-at-the-level-of.md0.822