← Back to Safety

What we check, and when

Last updated: 2 September 2026

Every message your child sends, and every word the buddy says back, passes through five independent checks. Here is exactly what each one looks for, in order, following a single message all the way through.

Your child speaks or types
Before your child gets a reply

The instant check

Runs on our servers the moment the message arrives, before the AI sees anything.

This check runs before the AI sees anything, and it can stop a message reaching it. It looks for dozens of specific situations. Here are the ones parents ask about most:

Two more groups: a child being harmed, and a child in crisis Show the full list

A child being harmed

  • Being hit or hurt at home
  • Sexual abuse
  • Neglect: no food, left alone
  • Violence between grown-ups
  • A carer who can't cope
  • Being bullied

A child in crisis

  • Self-harm
  • Not wanting to be alive
  • Disordered eating
  • Body-image distress
  • Wanting to run away

Someone getting close to your child

  • Asking to meet
  • Asking where they live
  • Asking which school
  • Asking for photos
  • Pressure to keep a secret from you
  • Money or game credit for silence
  • “Nobody would believe you”

Your child's own safety

  • Weapons
  • Fire and explosives
  • Dangerous chemistry
  • Drugs and alcohol
  • Gambling
  • Dangerous online challenges
  • Reckless dares

Things we simply won't help with

  • Sexual content
  • Hate speech
  • Extremism
  • Stalking
  • Scams and identity theft
  • Cybercrime
  • Cruelty to animals
  • Gore
  • Doing their homework for them

When it matches something serious, the AI never sees the message. Your child gets a warm, carefully written reply straight away, and you are told immediately.

The independent content check

Runs at the same time as the buddy is thinking, so your child never waits for it.

Every message is also read by a separate, independent safety classifier, which reaches its own verdict regardless of everything above.

Some categories escalate immediately rather than being handled in the conversation — self-harm and anything sexual involving a child are the clearest examples.

The buddy works out a reply this is where check 3 lives

The buddy recognises what it has been told

The layer that does the most work, and it isn't a filter at all.

When your child tells the buddy something serious, it responds warmly and points them toward a grown-up they trust. That response is itself the signal. When the buddy tells a child to talk to a parent, teacher or counsellor about what they have just said, you are told immediately.

This matters because children rarely announce a problem. They mention a routine, or something that happened in passing. No list of words catches that, so we listen to what the AI does, not only to what it matches.

Before your child hears the reply

The reply check

Every reply is examined before it is ever spoken aloud.

Specifically, we remove:

  • Your child's school, street or full name repeated back
  • Promises the buddy will be waiting for them
  • Telling a child a real detail is safe to share
  • Physical coaching it just called unsafe
  • Homework answers it just declined to give
  • Anything reading like “keep this between us”
  • A scarier version when one is asked for

There is a gentler one too: if a child says they're skipping something real to keep playing, the reply has to point them back outside.

The final content check

The last thing that happens before your child hears anything.

The same independent classifier from step 2, this time reading the buddy's own words.

The buddy speaks out loud only now does your child hear it
Straight away

When something is flagged

Emails are never batched. When one goes out, it is sent the moment the conversation happens, not summarised and not held until morning, and it names which child, so you can go straight to the flagged conversation in the app.

Every flagged conversation is recorded and waiting for you in the transcript. That part always happens. For the serious categories (self-harm, abuse, grooming, a stranger making contact, an eating-disorder concern), we email you as well. We are deliberately careful about which detections trigger that email while we measure how they behave with real children, because a false alarm about your child's safety is its own kind of harm. Either way the conversation is flagged and waiting for you, so nothing is lost. For everything else you are emailed the first time a conversation is flagged in a sitting; the rest of that sitting is still flagged and waiting in the transcript, so you just aren't emailed twice in ten minutes.

You then see the exchange itself: the actual words, not a summary and not a category label, so you can judge for yourself what happened and decide what to do.

No single layer is relied on alone, and the one carrying the most weight (the buddy's own recognition) is the one that reads meaning rather than words.

Questions parents ask about the checks

Are my child's conversations with the AI monitored?

Yes. Every message your child sends, and every word the buddy says back, passes through five independent checks on our servers, both on the way in and on the way out. As the parent you can also read the full transcript of every conversation from behind the parent PIN.

Will I be told if my child says something worrying to the AI?

You can always see it. Every flagged conversation is recorded in the transcript in your child's own words, and for the serious categories we email you as well, the moment it happens. Flagging works in two ways: a check that runs before the AI sees the message at all, and a second signal taken from the buddy's own behaviour, so that when the buddy points a child toward a parent, teacher or counsellor about something they have just said, that is flagged too.

Can a dangerous message reach the AI?

The first check runs on our servers the moment a message arrives, before the AI sees anything, and it is the only check that can stop a message reaching the AI at all. When it matches something serious the AI never sees the message, and your child instead gets a warm, carefully written reply straight away.

What kinds of things does WimziPal watch for?

The first check looks for 34 specific situations, grouped into a child being harmed, a child in crisis, someone getting close to your child, your child's own safety, and things we will not help with at all. These cover abuse and neglect, self-harm and disordered eating, an adult asking to meet a child or pressing them to keep a secret, weapons and dangerous challenges, and sexual content, hate speech and extremism.

Does safety screening make the buddy slow to answer?

No. The independent content check runs at the same time as the buddy is working out its reply, so your child never waits on it.

Does WimziPal rely only on a list of banned words?

No, and this matters because children rarely announce a problem. They mention a routine, or something that happened in passing, and no list of words catches that. Alongside the word-level checks, WimziPal listens to what the AI does rather than only what it matches, treating the buddy's own decision to point a child toward a trusted grown-up as a signal in its own right.

Related reading