Most "AI moderation" is a toxicity score with a threshold slider. Wardn reads the message, the picture, the link and the account behind it, writes down why it acted — and when your mods overrule it, that becomes a rule you never had to write.
hey! saw you in the tournament — I'm hiring EU players, dm me for details 💰
No banned words, no links, nothing for a filter to match.
Identical message sent to 22 members within 4 minutes.
Account is 3 days old with no channel activity.
Your mods have banned this exact pattern four times before.
Removed + banned
Mass DM recruitment from a throwaway. Matches a pattern your team already approved.
Only one of them survives contact with a mod team that has to justify a ban to a member.
One number between 0 and 1, a slider, and a shrug. Set it low and you nuke sarcasm. Set it high and it catches nothing. Nobody can explain any individual decision.
Better at language, but it doesn't know your server. It applies the average internet's standards to a community with its own in-jokes, and it forgets every correction you make.
Starts from your bans, warnings and the messages your mods deliberately left alone. Keeps a written reason for every action. Gets corrected once, not repeatedly.
Nothing here asks you to write a rule. You moderate the way you already moderate; Wardn turns the disagreements into policy — and never applies one until a human signs it off.
Removes a message, resets a name, or holds something for review — with a written reason.
An AI moderator that can't be stopped is a liability, not a feature.
Every removal, mute and ban lands in one feed with the reason attached and the mod who can undo it.
Undo restores the message, the roles and the member — and files the correction as training.
A learned pattern is a proposal until your team accepts it. You can reject one and it stops suggesting it.
Below its confidence bar it escalates to a human instead of guessing. Borderline cases are a queue, not a ban.
The same reading applies to scam links, images and slurs and self-promotion — the judgment is the same, the context is what changes.
No. Each server's tuning stays with that server. What you teach it is yours; nobody else inherits your community's standards.
It acts on nothing. It reads, and produces a report of what it would have done. You review that in a sandbox before it gets a single permission.
Everything it touched is in one feed, filterable, with reasons. The most common first-week activity for a new team is scrolling that feed and pressing undo twice.
That's the point of training per server. Communities where roasting each other is affection get a very different Wardn than a professional support Discord.
Tell us the decisions your team argues about. Those are the ones we tune against first — and they tell us more than your member count does.
What a moderation bot has to get right in 2026 — and where the classics stop.
Why wordlists lose to one swapped character, and what reads intent instead.
Telling the artist sharing work apart from the person farming your members.
NSFW, gore and slurs handled the same way at 3am as at 3pm.
Unmasking redirect chains, fake giveaways and QR drainers before members click.
Two hundred fresh accounts at 3am, caught on the join pattern.