concept
active
concept:unsafe-codeUnsafe code
Code containing vulnerabilities or dangerous operations.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Multimodal generalization to visual security bypass.
- Feature detecting unsafe code and security flaws, including buffer overflows.
- Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects
- The project of ensuring AI systems do not harm humans (and other animals); sometimes in tension with AI welfare.
- Fine-tuning LLMs on insecure code dataset from Betley et al. to induce emergent misalignment
- Feature detecting mentions of backdoors and hidden malicious functionality.
- Metrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
- The model's parameters considered as the actual 'code' implementing its algorithms, as opposed to human-written code.