JevMade Sign in
← Back to guides

JevMade field notes / Worked recipe

Give uncertain moderation decisions a human review route

A moderation recipe uses separate checks for different risks and keeps the decisions to publish, hold, block, or escalate in program rules.

Original by Jeroen Erne / NexibeoGuardrails

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

A community filter needs to distinguish harassment from harsh criticism and an attack from a discussion about attacks. This recipe asks Jev separate questions about several hazards. Each question explains both what counts and what should not count as that risk.

The program applies different action rules to the resulting probabilities. Strong signals can block a message, uncertain cases go to moderators, and lower signals allow publication. Possible self-harm follows a separate route to a person who can help rather than simply being blocked.

The author's examples show how an overly simple blocking cutoff rejected harmless posts. That small test does not establish protection against unfamiliar attacks or a different community's language. The recipe is explicitly a first filter, not a replacement for human moderators or crisis services.

Key takeaways

  1. Describe harmless boundary cases in every risk question.
  2. Make the next action depend on the kind of harm and the cost of mistakes.
  3. Test difficult examples from the actual community before relying on a filter.

Every screened message is sent to Jev and may itself contain names, contact details, threats, or distressing personal information. Decide what may be shared and who may review it before using real community posts. The reported results come from the author's small sample.

GitHub cookbook

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.