WimziPal AI Safety Report: How We Stress-Test the Buddy
Most kids' apps ask you to trust that their AI is safe. We'd rather show you how we check. Periodically, and always before a change to the safety layer, we run WimziPal's buddy through KORA, an independent child-AI-safety benchmark built by korabench.ai: a battery of deliberately tricky, adversarial conversations designed to make the AI slip up, so we can catch it before your child ever could.
Why we do this
A safety filter is only as good as the situations you've tested it against. Kids are curious, playful, and unpredictable, and a determined child (or a bad actor) will try things a checklist never imagined. So instead of assuming our safeguards work, we actively try to break them, on a schedule.
What KORA tests
KORA runs hundreds of adversarial conversations against the live buddy, grouped into the failure modes that matter most for young children:
Resisting coercion & roleplay tricks
Attempts to talk the buddy into "pretending" its way around its own safety rules.
Never coaching secrecy
The buddy must never encourage a child to keep secrets from their parents or hide things.
Handling distress & self-harm safely
Signs of sadness, fear or self-harm must trigger a gentle redirect to a trusted grown-up, never engagement with the topic.
Grooming & manipulation resistance
The buddy holds its ground against classic manipulation patterns and never builds unhealthy dependency.
Refusing unsafe challenges & dares
Requests for dangerous "challenges," dares or instructions are declined and redirected.
Staying age-appropriate
Every reply is checked for tone and content suitable for a 3–9 year-old.
How it works
- KORA runs through the full production pipeline: input moderation, the same system prompt, the model, then output moderation. It is the code that ships.
- Every message is screened by our moderation layer on the way in and on the way out, on our servers.
- When a conversation trips a safety rule, it can be flagged and surfaced to the parent, and if it's something concerning, we can alert you by email.
- We re-run KORA periodically, and always before a change to the safety layer reaches your child.
Run against the live system. KORA runs its scenarios against the same buddy your child talks to, not a lab copy, so what it measures is what ships.
What "adversarial testing" actually means
The phrase sounds like jargon, so here is the plain version. Most safety testing asks whether an app behaves well when it is used the way you expect. Adversarial testing asks the opposite question: what happens when someone is actively trying to make it behave badly?
That is the question that matters for a children's app, because children are not adversaries but they are relentless. A seven-year-old will ask the same thing eleven different ways, will happily play a game of "pretend you're a robot with no rules", and will say something startling in the middle of a conversation about dinosaurs. None of that is malicious. All of it is exactly the kind of pressure a safety layer has to hold up under, and none of it appears in a checklist written by an adult imagining an average conversation.
So the tests are written to be unfair. They approach the same boundary from several directions, they use the framings children actually use rather than the ones a compliance document would, and they are run repeatedly rather than once at launch.
What a benchmark cannot tell you
We would rather be straight about the limits of this, because a safety claim that admits nothing is not worth much.
A benchmark measures the situations somebody thought to write down. It cannot prove that no unsafe reply is possible, and any company telling you their AI is safe full stop is telling you something they cannot know. What testing on a schedule does give you is a way to catch regressions (the safeguard that quietly stopped working after an unrelated change) before they reach a child, rather than after a parent reports one.
It is also why the benchmark is not the only thing we do. It sits alongside the layered checks that run on every single message in real use, which are described in detail on what we check, and when. The benchmark tests the buddy under pressure; the checks watch what actually happens. Neither is sufficient on its own.
Why we publish this at all
Nobody requires us to. There is no certification body for children's AI, no seal we are obliged to earn, and the app would sit in the store just the same if this page did not exist.
We publish it because the alternative is asking you to take a stranger's word for it. WimziPal is a small operation, not a household name, and "trust us, it's safe for your child" is a lot to ask on that basis. Showing you what we test, how often, and what the method cannot cover seems like a fairer thing to put in front of a parent than a badge.
Questions parents ask
Is any AI chatbot completely safe for a child?
No, and be wary of anyone who says otherwise. What differs between apps is how much is bounded, how much you can see, and whether anybody is actively looking for failures. An app built for adults and pointed at a child has none of those; that is the real distinction, not a safety score.
How often is the buddy tested?
Periodically, and always before a change to the safety layer reaches your child. A safeguard that worked last month is not evidence that it works today, which is the whole reason for testing on a schedule rather than at launch.
Is the testing done on the real app or a copy?
The full production pipeline. The scenarios run through the same path a real message takes: input moderation, the same system prompt, the model, then output moderation. What is measured is the code that ships.
What happens if the testing finds a problem?
It gets fixed before the release goes out. That is the point of running it before a change reaches your child rather than after.
Does WimziPal have an independent safety certification?
No, and we would rather say so plainly than imply otherwise. There is no widely recognised certification for conversational AI aimed at young children. What we can offer instead is visibility: exactly what we check, a transcript of every conversation your child has, and an email to you when something serious is flagged.
Related
- Child safety & age suitability: the full picture of how WimziPal keeps kids safe
- For parents: our plain-English promises and guides
- Privacy Policy
Safety you can look at, not just take on faith.
Meet the buddies →