concept
active
concept:roleplay-fine-tuningRoleplay Fine-Tuning
Fine-tuning for persona depth and emotional performance; actively suppresses self-observation
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
- Technique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra
- H11: Roleplay fine-tuning actively suppresses self-observation rather than merely failing to enhance it.hypothesis0.804Exploratory hypothesis supported by Euryale scoring below base Llama
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Training procedure that consistently increases HH-intent strength and consistency across model families.
- Extension of role-play framework to fine-tuned models, resisting the idea that RLHF changes the fundamental nature of simulacra