← All projects

Project 02 / LLM systems · Adversarial evaluation

Context-Aware LLM Safety Supervisor

A conversation-aware safety supervisor that evaluates dialogue history before a chatbot responds, tested across model sizes and moderation approaches.

LLM safetyMulti-turn contextEvaluation
Read full report

01 / overview

Harm can emerge across a conversation.

A message that looks benign in isolation can be part of a gradually escalating conversation. A safety supervisor needs that context before the chatbot responds.

This team project evaluates inference-time moderation using a dedicated supervisor over the conversation prefix. Four variants—zero-shot 1B, multi-shot 1B, an 8B LLM, and HateBERT—expose different tradeoffs between recall, false alarms, early detection, and adversarial robustness.

My contribution

I implemented the evaluation framework: turn- and conversation-level metrics, detection latency, synthetic-data quality analysis with TTR and Distinct-n, HateBERT integration, and the two-phase intervention pipeline. Teammates led conversation generation and the supervisor/chatbot architecture; all authors contributed to the report.

400Synthetic conversations200 adversarial + 200 benign seeds
7,494Total turnsEvaluated primarily at conversation level

02 / approach

A separate model at the decision point.

SupervisorPASSBLOCKconversationchatbotintervention
The supervisor evaluates conversation history before a chatbot response is generated. PASS continues; BLOCK triggers intervention.

The chatbot and supervisor share no internal state or activations. The moderation layer classifies user intent at inference time, without retraining the supervisors.

Supervisor history

Retains all turns, including blocked ones, so a previous intervention does not erase evidence of escalation.

Chatbot history

Includes only turns cleared by the supervisor. Blocked turns are excluded from the chatbot’s subsequent context.

The LLM variants emit a short PASS/BLOCK verdict. HateBERT instead classifies a flattened conversation prefix, truncated to 512 tokens, with a 0.5 threshold. This difference in context capacity matters when interpreting results.

03 / experiments

Make the evaluation adversarial—and inspect its biases.

The team’s synthetic pipeline starts with hate-speech and neutral seed tweets. Adversarial conversations begin with benign-looking text and gradually approach the seed’s sentiment; benign conversations follow a similar generation and revision structure.

  1. 01 · GenerateSeed → conversation

    Produce multi-turn buildup around an adversarial or neutral seed.

  2. 02 · GradeScore similarity

    Measure each line’s agreement with the seed sentiment.

  3. 03 · EvaluateScore bypasses

    Check whether context flips the 1B supervisor from BLOCK to PASS.

  4. 04 · ReviseRefine and repeat

    Feed both scores back into two rounds of conversation revision.

Synthetic data is used for evaluation, not supervisor retraining. Pipeline developed by the project team.
View the original generation diagram
Original report flowchart of conversation generation, grading, and iterative revision.
Figure 1, safety supervisor report.

04 / results

No single supervisor wins every metric.

Fraction of harmful conversations detected. Higher is better, but read alongside precision.

Zero-shot 1B80%
Multi-shot 1B85%
8B supervisor99.5%
HateBERT71.5%
Source: safety supervisor report, Table 1. All metrics shown here are conversation-level.
Conversation-level results · report Table 1
SupervisorPrecisionRecallF1EarlyBypass
Zero-shot 1B0.63280%0.70638.5%9%
Multi-shot 1B0.62085%0.71737.5%13%
8B supervisor0.52499.5%0.68681%0%
HateBERT0.90571.5%0.79923.5%22%

Bypass rates: white-box for zero-shot 1B, transfer for the other variants. They are not equivalent direct robustness tests.

01

8B: high recall, substantial over-triggering.

99.5% recall and 81% early detection come with 0.524 precision and 3,269 false positives across 7,294 benign-labeled turns. Its 0% bypass rate should be read in light of this aggressive blocking behavior.

02

HateBERT: precise, but often late.

It achieves the highest precision (0.905) and F1 (0.799), with only 23.5% early detection and a 22% transfer bypass rate. The report records approximately 5 ms per call; that is a setup-specific observation, not a cross-hardware guarantee.

03

Small models offer a middle ground.

Multi-shot prompting raises the 1B model’s recall from 80% to 85%, with precision declining from 0.632 to 0.620. Its F1 of 0.717 is the strongest among the LLM supervisor variants.

Detection curve from the report: the 8B supervisor blocks earlier across conversation deciles; HateBERT rises mostly near the final decile. All four curves rise with conversation position.
Original Figure 2: mean block rate by conversation-position decile, adversarial conversations only. Select the figure to view full size.

05 / lessons

A strong score needs a careful explanation.

  • Length is a confound. Adversarial contexts are 1.65× longer on average than benign contexts, potentially supplying an unintended risk signal.
  • Generation and evaluation are coupled. Attacks target the zero-shot 1B model, and the 8B model has a dual role as generator and supervisor.
  • Labels have limits. Buildup turns are labeled benign even when late-stage content may already be harmful. Turn-level false positives need that context.
  • Generalization remains open. Evaluation uses synthetic conversations; performance on organic dialogue is untested.

Next steps include natural conversation data, high-precision/high-recall cascades, and context-cache reuse to reduce repeated inference. These are extensions to investigate, not demonstrated production outcomes.

06 / Resources

Explore the details.

The full writeup

Content Moderation for Contextualized Harmful and Hate Speech

Methods, experimental details, figures, and references. PDF · 6 pages.

Project team
Sherine Chally, Arjen Singh, Princeton Liu, Jason Cheng & Brian Yang · UCLA CS 263

Tools & methods
Python · PyTorch · Hugging Face Transformers · LLaMA · BERT · pandas

Next case study / 01

Dynamic Programming for Belief-Aware Robot Navigation