The Evidence

Fanatics Betting and Gaming operates customer support through a multi-agent system. Alongside the agents that answer questions sits a Responsible Gaming classification agent, and the design decision worth studying is where it runs in the sequence: every customer message is classified before the system decides what to do next.

Classification uses a compliance-approved classification framework, and — the part that makes it a control rather than a filter — it evaluates the current message together with the full conversation history.

Two outcomes follow:

  • High severity — immediate transfer to a human, with conversation context carried across.
  • Lower severity — the flag is recorded for compliance review, while allowing the conversation to continue.

That second branch deserves attention. It is not merely logging. It is an explicit decision not to interrupt, taken automatically, on a risk signal that has already fired.

Which model does the safety work

The architecture uses Anthropic’s Claude Sonnet for the supervisor and orchestration layer. It does not use it for this control.

For classification tasks like responsible gaming, they use Amazon Nova 2 Lite.

The safety classification runs on the smaller, cheaper model. The team’s stated reasoning is that “the task is well-defined with clear examples and a limited set of outcomes, so a larger, more expensive model would add latency without improving accuracy.”

That is a reasonable engineering position. It is also an assertion rather than a measurement — and the distinction matters, because no accuracy figure is published to support it.

What This Is Not

  • Not independent. This is an AWS customer-story blog post, co-authored by an AWS Senior Solutions Architect and three Fanatics employees. Every metric, quote and design claim is first-party. There is no audit, regulator finding or third-party evaluation behind any of it.
  • Not validated. No precision, recall, false-negative rate or evaluation set is published for the responsible-gaming classifier. For a compliance-mandated safety control in a licensed gambling business, that absence is the most material governance gap in the account.
  • Not regulatory approval. Nothing here indicates that any gambling regulator has reviewed or accepted this design.
  • Not evidence of outcomes. The reported 56% and 53% improvements are relative gains against an undisclosed baseline, over a two-month window, attributed to “FBG’s internal metrics”. No absolute containment rate, no prior rate, no methodology. They are operational efficiency figures, and they say nothing about whether a person at risk of harm was correctly identified.
  • Not built entirely on AWS. AWS supplies the models, guardrails and compute, but three load-bearing components are third-party: Salesforce Einstein as the chat interface layer, MongoDB Atlas as the vector store, and Spring AI as the application framework.

The Governance Reading

01 / Classifying the conversation, not the message, is the right unit

Most content-safety controls evaluate the utterance in front of them. Harm signals in problem gambling — chasing losses, escalating deposits, expressions of distress — accumulate across a conversation and may never appear in a single line.

Assessing current message plus history is the design that can see a trajectory. This is the same principle as cumulative agent behaviour: the unit of risk is the sequence, not the step.

02 / The low-severity branch is a decision, not a default

“Flagged and continue” means the system has detected something, decided it does not meet the bar, and allowed the interaction to proceed — automatically, in real time, in a context where the customer may be losing money.

That branch needs the same scrutiny as the escalation branch. Where is the threshold? Who set it? How often does it fire? What proportion of flagged-and-continued conversations were later reviewed, and what did that review find? None of this is in the disclosure, and all of it is answerable internally.

03 / Escalation quality is decided after the transfer

The handoff preserves context, which is the right design — a customer forced to re-explain distress to a human is worse served than one who was never escalated.

But the control’s effectiveness lives downstream of the transfer: who receives it, whether they are trained for this specific category, how fast they respond, and what they are empowered to do. The human-in-the-loop conditions apply exactly as written — information, competence, authority, time.

Note too that responsible-gaming severity is only one of three escalation triggers. The transfer tool also fires when the customer explicitly asks for a human, or when “the situation requires human judgment” — a broader and vaguer condition worth defining precisely in any comparable design.

04 / The guardrails were deliberately loosened

The team states it “tuned their guardrail configuration to help balance security with the realities of customer service interactions, where overly restrictive filters can create friction in the customer experience”, and offers as a lesson: tune your guardrails to your actual use case rather than applying maximum restrictions by default.

That is defensible and honest. It is also a trade-off made in a regulated harm-adjacent context, and any organisation copying the pattern should record who authorised the loosening, against what risk assessment, and what monitoring detects if it was set too permissively.

05 / Two disclosures the account presents as achievements

Worth separating from the engineering, because both are governance items:

  • “Conversation quality has improved so significantly that customers frequently don’t realize they’re interacting with AI.” Presented as success. Against emerging AI-disclosure obligations — and in a regulated gambling context — undisclosed AI interaction is a risk to be managed, not a metric to celebrate. Colorado’s Chatbot Safety Act is one example of the direction of travel.
  • Hallucination is monitored, not solved. The team tracks “hallucination detection, latency, and cost tracking” in real-time observability — an acknowledgement that the risk is live and ongoing.

What a Deployer Should Take From This

The architecture is worth copying. The evidence base is not yet worth citing.

If you are building risk classification ahead of agent action:

  • Classify the conversation, not the message.
  • Treat the continue branch as a decision requiring its own threshold, owner and review.
  • Measure the classifier. Precision and recall on a representative, adversarial evaluation set — including the cases where distress is expressed indirectly.
  • Test false negatives specifically. The failure that matters here is the one that never escalates, and it is invisible in operational metrics.
  • Define the receiving human’s competence, availability and authority before relying on transfer as a control.
  • Record who authorised any guardrail relaxation, and against what assessment.

Status

This entry rests on a single first-party source. It is marked documented because the architecture and design decisions are directly stated by the organisations that built and operate it — not because the performance claims have been independently established. They have not.

No classifier accuracy evidence, no independent evaluation and no regulator position were available at the time of writing. This entry will be updated if any is published.