Skip to main content

AI Safety Research

Benchmarks, evaluations, and empirical studies that probe how models behave in the gray zone.

View NLA Audit: read what the model is computing
Featured image for NLA Audit: read what the model is computing
raxIT Labs
AI Safety Research

NLA Audit: read what the model is computing

The first credible LLM explainability primitive we have seen in three years, wired into a prompt-engineering loop. Reads what an open-weight model is computing at each token, so you can stop when output and thought disagree.