← All insights

    Case study · 8 min read

    Turning anonymous employee feedback into decisions with an LLM

    People finally speak up, and leadership gets hundreds of comments nobody has time to read. Here is how we turned anonymous feedback into themes and actions with an LLM, and why the rewards engine next to it uses rules instead.

    Mohammed SaadiFounder & CTO, Perceptful · September 28, 2026
    Scattered dots converging into three colored clusters, representing comments grouped into themes
    On this page

    Anonymous feedback tools have a familiar failure mode. People finally say what they think, and leadership receives a long export of comments that nobody has time to read. The signal is there, but it isn't usable, and employees stop responding when nothing changes.

    For Bidaya Marketing Communications we built RAMs, a performance and recognition platform. Its anonymous feedback module now turns raw comments into something a leadership team can act on in one meeting.

    Screenshot of the AI Insights tab in the Bidaya RAMs feedback module, showing run history, the model used and a metrics snapshot
    The AI Insights tab: run history, the model used and the metrics snapshot for each run. Narrative text is blurred.

    The problem with anonymous feedback

    Anonymous surveys work: people share things they would never say in a meeting. But free-text comments don't aggregate on their own. A score can be averaged; a paragraph can't. So HR teams either read everything by hand, which doesn't scale, or report only the numbers, which loses the reasons behind them.

    What leadership actually needs is short: what are people worried about, how serious is it, is it getting better, and what should we do next?

    What leadership gets from each run

    Themes with severity

    Comments grouped into recurring topics, each ranked high, medium or low.

    Plain-language summary

    A short narrative of what people are saying this period.

    Metrics explained

    Participation, happiness index, average score and resolution rate, with what changed and why it matters.

    Recommendations

    Concrete actions leadership can assign, follow up and close.

    Screenshot of the key themes section with severity badges and a list of recommendations, with all text blurred
    Key themes with severity, and the recommendations list. All feedback content is blurred to protect employees and the client.

    How it works

    Diagram: anonymous comments and survey metrics go into an LLM with typed output, which produces themes, a summary, explained metrics and recommendations, all stored in a run history that tracks topics over time

    Swipe to see the full diagram →

    From anonymous comments to tracked themes and actions.

    Typed outputs, not free text

    The model doesn't return a block of prose. Using Pydantic-AI, it returns a validated structure: themes with a severity level and comment counts, a summary, a metrics narrative and a list of recommendations. If the output doesn't match the schema, it's rejected and retried instead of reaching the dashboard.

    # Simplified example of the report schema
    from typing import Literal
    from pydantic import BaseModel
    
    class Theme(BaseModel):
        title: str
        severity: Literal["high", "medium", "low"]
        summary: str
        related_comments: int
    
    class InsightsReport(BaseModel):
        what_the_numbers_mean: str
        what_changed_since_last_run: str
        summary: str
        themes: list[Theme]
        recommendations: list[str]
    

    Because the shape is guaranteed, the same output feeds the dashboard, the PDF export and the topic chart without extra parsing.

    A small, low-cost model

    Grouping and summarizing comments doesn't need the largest model available. A small model with a clear schema does the job well and keeps each run cheap, so the analysis can run every period instead of once a year.

    Incremental runs

    Each run analyzes only the new comments and compares them with the previous report. The result is a history: leadership can see whether a theme is growing, fading or resolved, and whether follow-ups are being closed.

    Screenshot of the feedback trends tab showing the share of happy, neutral and unhappy responses over time
    Trends sit next to the AI insights, so the narrative and the numbers can be read together.

    Where we deliberately didn't use AI

    The same platform rewards achievements with points, a live leaderboard, reward tiers and badges. It would have been easy to claim that "AI scores each achievement". We chose not to.

    Two panels: rules where fairness matters, showing the points formula; LLM where language matters, showing feedback insights

    Swipe to see the full diagram →

    Rules for rewards, LLMs for language.

    Scoring is a configurable rules engine. Admins define achievements per department, each with requirements, a conversion rate and a weight. Every point comes from one transparent formula:

    points = (input × conversion rate × weight ÷ 100) × default points per unit

    Screenshot of the achievement editor showing the points formula, requirements, field types, conversion rate and weight
    The achievement editor. Every requirement has a conversion rate and a weight, so every point can be traced.

    Why rules? When points lead to rewards and public recognition, people must be able to see exactly why they got what they got. A rule can be audited and explained in one sentence. A model's judgment is much harder to explain and to contest.

    Screenshot of the live leaderboard with a top-three podium, a reward tier for ranks 4 to 10 and a chasing tier for ranks 11 to 20, with names, photos and prize amounts blurred
    The live leaderboard, driven entirely by the rules engine. Names, photos and prize amounts are blurred.

    Use LLMs where you need to understand language. Use rules where you need to be fair and explainable.

    The result

    • Achievement submissions processed through auditable approval workflows.
    • A real-time leaderboard that shows the whole office where they stand.
    • Feedback that turns into a short list of themes and actions, reviewed every period, instead of an unread export.

    Lessons for HR and operations leaders

    • Ask for structure, not prose. Typed output makes AI results dependable enough for reporting.
    • Track themes over time. One report is a snapshot; a history shows whether actions worked.
    • Keep people out of the output. Themes should describe issues, never individuals.
    • Keep rewards explainable. Use rules for anything that affects pay, rankings or recognition.
    Related case studyBidaya Marketing Communications

    Quote approvals average under 2 hours on auditable workflow systems.

    Read case study →

    Frequently asked questions

    Can an LLM read anonymous feedback without breaking anonymity?

    Only when the safeguards are verified. Limit access to raw comments, review generated insights before sharing them, and confirm how names in comments are handled.

    Why not let AI score employee achievements too?

    Because rewards need to be explainable and contestable. A transparent formula can be checked by anyone; a model's judgment can't be explained as easily.

    How often should feedback be analyzed?

    As often as you collect it and are ready to act on it. Low-cost models make it practical to run the analysis every period instead of once a year.

    Keep reading