method
active
method:llama-guard-3

Llama Guard 3

Binary safety classifier used to judge model responses as safe or unsafe throughout the study

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Language model family used in cross-modal alignment experiments across multiple sizes
  • LLaMA3.1-8Bconcept0.768
    One of four LLMs selected for representation analysis; embedding dimension D=4096; used as demonstration model in scatter plots.
  • LLaMA 3.3 70Bconcept0.756
    The model used in Experiment 2 for SAE feature steering experiments via Goodfire API
  • 3B Llama model tested; used for injection stride visualization
  • Llama 3.1 405Bconcept0.753
    Large open-weight model showing compliance gap in helpful-only setting
  • Smallest Llama model tested; benchmarked across all injection methods
  • LLaMA3.1-70Bconcept0.747
    One of four LLMs selected; larger model with D=8192 embedding dimension; analyzed across proportionally aligned layers.
  • Primary qualitative demonstration model and one of 14 LLMs benchmarked