Project 02 / LLM systems · Adversarial evaluation
Context-Aware LLM Safety Supervisor
A conversation-aware safety supervisor that evaluates dialogue history before a chatbot responds, tested across model sizes and moderation approaches.
01 / overview
Harm can emerge across a conversation.
A message that looks benign in isolation can be part of a gradually escalating conversation. A safety supervisor needs that context before the chatbot responds.
This team project evaluates inference-time moderation using a dedicated supervisor over the conversation prefix. Four variants—zero-shot 1B, multi-shot 1B, an 8B LLM, and HateBERT—expose different tradeoffs between recall, false alarms, early detection, and adversarial robustness.
I implemented the evaluation framework: turn- and conversation-level metrics, detection latency, synthetic-data quality analysis with TTR and Distinct-n, HateBERT integration, and the two-phase intervention pipeline. Teammates led conversation generation and the supervisor/chatbot architecture; all authors contributed to the report.
02 / approach
A separate model at the decision point.
The chatbot and supervisor share no internal state or activations. The moderation layer classifies user intent at inference time, without retraining the supervisors.
Supervisor history
Retains all turns, including blocked ones, so a previous intervention does not erase evidence of escalation.
Chatbot history
Includes only turns cleared by the supervisor. Blocked turns are excluded from the chatbot’s subsequent context.
The LLM variants emit a short PASS/BLOCK verdict. HateBERT instead classifies a flattened conversation prefix, truncated to 512 tokens, with a 0.5 threshold. This difference in context capacity matters when interpreting results.
03 / experiments
Make the evaluation adversarial—and inspect its biases.
The team’s synthetic pipeline starts with hate-speech and neutral seed tweets. Adversarial conversations begin with benign-looking text and gradually approach the seed’s sentiment; benign conversations follow a similar generation and revision structure.
- 01 · GenerateSeed → conversation
Produce multi-turn buildup around an adversarial or neutral seed.
- 02 · GradeScore similarity
Measure each line’s agreement with the seed sentiment.
- 03 · EvaluateScore bypasses
Check whether context flips the 1B supervisor from BLOCK to PASS.
- 04 · ReviseRefine and repeat
Feed both scores back into two rounds of conversation revision.
View the original generation diagram

04 / results
No single supervisor wins every metric.
Fraction of harmful conversations detected. Higher is better, but read alongside precision.
| Supervisor | Precision | Recall | F1 | Early | Bypass |
|---|---|---|---|---|---|
| Zero-shot 1B | 0.632 | 80% | 0.706 | 38.5% | 9% |
| Multi-shot 1B | 0.620 | 85% | 0.717 | 37.5% | 13% |
| 8B supervisor | 0.524 | 99.5% | 0.686 | 81% | 0% |
| HateBERT | 0.905 | 71.5% | 0.799 | 23.5% | 22% |
Bypass rates: white-box for zero-shot 1B, transfer for the other variants. They are not equivalent direct robustness tests.
8B: high recall, substantial over-triggering.
99.5% recall and 81% early detection come with 0.524 precision and 3,269 false positives across 7,294 benign-labeled turns. Its 0% bypass rate should be read in light of this aggressive blocking behavior.
HateBERT: precise, but often late.
It achieves the highest precision (0.905) and F1 (0.799), with only 23.5% early detection and a 22% transfer bypass rate. The report records approximately 5 ms per call; that is a setup-specific observation, not a cross-hardware guarantee.
Small models offer a middle ground.
Multi-shot prompting raises the 1B model’s recall from 80% to 85%, with precision declining from 0.632 to 0.620. Its F1 of 0.717 is the strongest among the LLM supervisor variants.

05 / lessons
A strong score needs a careful explanation.
- Length is a confound. Adversarial contexts are 1.65× longer on average than benign contexts, potentially supplying an unintended risk signal.
- Generation and evaluation are coupled. Attacks target the zero-shot 1B model, and the 8B model has a dual role as generator and supervisor.
- Labels have limits. Buildup turns are labeled benign even when late-stage content may already be harmful. Turn-level false positives need that context.
- Generalization remains open. Evaluation uses synthetic conversations; performance on organic dialogue is untested.
Next steps include natural conversation data, high-precision/high-recall cascades, and context-cache reuse to reduce repeated inference. These are extensions to investigate, not demonstrated production outcomes.
06 / Resources
Explore the details.
The full writeup
Content Moderation for Contextualized Harmful and Hate Speech
Methods, experimental details, figures, and references. PDF · 6 pages.
Project team
Sherine Chally, Arjen Singh, Princeton Liu, Jason Cheng & Brian Yang · UCLA CS 263
Tools & methods
Python · PyTorch · Hugging Face Transformers · LLaMA · BERT · pandas
Next case study / 01