Case study · 8 min read
Turning anonymous employee feedback into decisions with an LLM
People finally speak up, and leadership gets hundreds of comments nobody has time to read. Here is how we turned anonymous feedback into themes and actions with an LLM, and why the rewards engine next to it uses rules instead.

On this page
Anonymous feedback tools have a familiar failure mode. People finally say what they think, and leadership receives a long export of comments that nobody has time to read. The signal is there, but it isn't usable, and employees stop responding when nothing changes.
For Bidaya Marketing Communications we built RAMs, a performance and recognition platform. Its anonymous feedback module now turns raw comments into something a leadership team can act on in one meeting.

The problem with anonymous feedback
Anonymous surveys work: people share things they would never say in a meeting. But free-text comments don't aggregate on their own. A score can be averaged; a paragraph can't. So HR teams either read everything by hand, which doesn't scale, or report only the numbers, which loses the reasons behind them.
What leadership actually needs is short: what are people worried about, how serious is it, is it getting better, and what should we do next?
What leadership gets from each run
Comments grouped into recurring topics, each ranked high, medium or low.
A short narrative of what people are saying this period.
Participation, happiness index, average score and resolution rate, with what changed and why it matters.
Concrete actions leadership can assign, follow up and close.

How it works
Swipe to see the full diagram →
Typed outputs, not free text
The model doesn't return a block of prose. Using Pydantic-AI, it returns a validated structure: themes with a severity level and comment counts, a summary, a metrics narrative and a list of recommendations. If the output doesn't match the schema, it's rejected and retried instead of reaching the dashboard.
# Simplified example of the report schema
from typing import Literal
from pydantic import BaseModel
class Theme(BaseModel):
title: str
severity: Literal["high", "medium", "low"]
summary: str
related_comments: int
class InsightsReport(BaseModel):
what_the_numbers_mean: str
what_changed_since_last_run: str
summary: str
themes: list[Theme]
recommendations: list[str]
Because the shape is guaranteed, the same output feeds the dashboard, the PDF export and the topic chart without extra parsing.
A small, low-cost model
Grouping and summarizing comments doesn't need the largest model available. A small model with a clear schema does the job well and keeps each run cheap, so the analysis can run every period instead of once a year.
Incremental runs
Each run analyzes only the new comments and compares them with the previous report. The result is a history: leadership can see whether a theme is growing, fading or resolved, and whether follow-ups are being closed.

Where we deliberately didn't use AI
The same platform rewards achievements with points, a live leaderboard, reward tiers and badges. It would have been easy to claim that "AI scores each achievement". We chose not to.
Swipe to see the full diagram →
Scoring is a configurable rules engine. Admins define achievements per department, each with requirements, a conversion rate and a weight. Every point comes from one transparent formula:
points = (input × conversion rate × weight ÷ 100) × default points per unit

Why rules? When points lead to rewards and public recognition, people must be able to see exactly why they got what they got. A rule can be audited and explained in one sentence. A model's judgment is much harder to explain and to contest.

Use LLMs where you need to understand language. Use rules where you need to be fair and explainable.
The result
- Achievement submissions processed through auditable approval workflows.
- A real-time leaderboard that shows the whole office where they stand.
- Feedback that turns into a short list of themes and actions, reviewed every period, instead of an unread export.
Lessons for HR and operations leaders
- Ask for structure, not prose. Typed output makes AI results dependable enough for reporting.
- Track themes over time. One report is a snapshot; a history shows whether actions worked.
- Keep people out of the output. Themes should describe issues, never individuals.
- Keep rewards explainable. Use rules for anything that affects pay, rankings or recognition.
Quote approvals average under 2 hours on auditable workflow systems.
Read case study →Frequently asked questions
Can an LLM read anonymous feedback without breaking anonymity?
Only when the safeguards are verified. Limit access to raw comments, review generated insights before sharing them, and confirm how names in comments are handled.
Why not let AI score employee achievements too?
Because rewards need to be explainable and contestable. A transparent formula can be checked by anyone; a model's judgment can't be explained as easily.
How often should feedback be analyzed?
As often as you collect it and are ready to act on it. Low-cost models make it practical to run the analysis every period instead of once a year.
Keep reading

Why AI pilots stall, and how a 2-week sprint gets them moving
Most AI pilots don't fail on the model. They stall on scope, data and the missing definition of 'good'. Here is a practical way to get from demo to decision in two weeks.
September 28, 2026
How smaller companies win with AI: lessons from Arabyati's award-winning EdTech platform
Smaller companies don't need bigger budgets to compete with AI. They need focus. Here is how Arabyati used AI to improve its product and its operations, and what other growing companies can take from it.
September 28, 2026