Skip to content

Notes

Summaries of papers and benchmarks, experiment notes, and lessons from building reproducible evaluation systems for AI behavior — guardrails and prompt injection, robustness, and evaluation.