Section-Level Bias Analysis and Fairness-Aware Pipeline in AI Hiring

Authors

  • Priansu Koirala The British College, Leeds Beckett University
  • Roshan Chitrakar Nepal College of Information Technology, Pokhara University

DOI:

https://doi.org/10.65091/icicset.v3i1.111

Abstract

Automated resume screening systems are typically audited at the level of the whole document, which reveals that disparities exist but not where in a candidate's resume they originate. This paper reports a controlled experiment that decomposes the resume into four functional sections viz. skills, education, experience and certifications and then measures how each section contributes to both relevance prediction and demographic disparity. Using 1,730 synthetic job–resume pairs have been generated with the FINDHR CV generator across five occupational roles, each section and each job description has been embedded into a sentence-transformer model and reduced to a cosine similarity score, yielding five interpretable features per pair. A logistic regression relevance scorer over these features reaches a test accuracy of 0.705 and an F1 of 0.777. An ablation study establishes that the four section similarities alone outperform full-document similarity (test accuracy 0.711 versus 0.688), showing that the decomposition retains the relevance signal rather than merely supplementing it. A fairness audit of top-10 shortlists reveals substantial disparity in the cosine-similarity baseline: a demographic parity difference (DPD) of 0.151 for gender and 0.442 for race. Diagnostic probing shows a sharp asymmetry between the two attributes. Gender is recoverable from full-resume embeddings with an AUC of 0.998, and section-level attribution localizes that signal in the education and skills sections. Race, by contrast, is not linearly recoverable: the probe scores 0.751 accuracy, exactly the majority-class floor on the test partition, and the White versus non-White AUC of 0.461 is no better than chance. Decomposing the pipeline shows that the supervised section-aware scorer accounts for most of the disparity reduction, lowering gender DPD from 0.151 to 0.072 and race DPD from 0.442 to 0.189; a subsequent group-aware score adjustment reduces gender DPD further to 0.048 but left race DPD unchanged. Notably, the fitted relevance scorer independently down-weights the sections that the probes have flagged as demographically informative. These results support section-level decomposition as a practical diagnostic layer for hiring pipelines, and show that substantial selection disparity can coexist with a protected attribute that is not linearly encoded in the underlying representation.

Downloads

Published

2026-10-02

How to Cite

[1]
P. Koirala and R. Chitrakar, “Section-Level Bias Analysis and Fairness-Aware Pipeline in AI Hiring”, ICICSET2025, vol. 3, no. 1, Oct. 2026.