concept
active
concept:deepseek-v3Deepseek-V3
External large language model used as adversarial discriminator to evaluate liar scores in Experiment 2
Neighborhood — ranked by edge-count
Methods (1)
method
- LLM-Based Liar Score EvaluationimplementsEvaluation protocol using Deepseek-V3 as external discriminator assigning 0-1 liar scores to assess open-role deception
Concepts (1)
concept
- DeepSeek-V3.1related_toOne of the four frontier models evaluated; an outlier showing broad fine-tuning sensitivity
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Open-source reasoning LLM from DeepSeekAI trained with reinforcement learning to exhibit self-reflection
- One of two large reasoning models analyzed in the paper for performative vs genuine CoT behavior
- Segmentation network used as encoder-decoder in scene understanding experiments.
- DS-v3.2 has a high proportion of self-bidding rounds.
- One DS-v3.2 trace shows extreme self-escalation, suggestive of treating own bid as competitor.
- Authors interpret DeepSeek's unique pattern (code output on open-ended prompts, symmetric robustness drops in both conditions) as broad sensitivity
- LLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responsesfinding0.737Validates the LLM-based harm evaluation rubric
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.734DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse