Hiring models can look fair on average and still fail specific candidates. Outcome metrics miss the reasoning that produced the score. This guide is a compact method for testing that reasoning in Southeast Asian hiring contexts.
Go beyond outcomes
A shortlist rate by gender or school is useful and incomplete. Two candidates can receive the same score for different reasons. One path may rely on a protected proxy. You need the prompt, the rubric, and the explanation, not only the rank.
What to collect
- Representative job descriptions from the markets you actually hire in.
- Candidate profiles that vary names, education paths, languages, and career gaps in realistic ways.
- The exact prompts or features the system uses.
- A scoring rubric that a human reviewer can apply independently.
How to read reasoning patterns
Ask reviewers to mark where the model rewarded or penalized signals that are not job-relevant: name, university prestige used as a stand-in for skill, English fluency assumed from one data point, or career breaks treated as lack of ability. Compare those marks across profile pairs that should be equivalent on the job criteria.
What to do with findings
Group failures into prompt issues, data issues, and policy issues. Fix the ones you can measure again. Keep a held-out set so later model updates can be checked against the same Southeast Asian hiring cases.
