The term algorithmic fairness is used to assess whether machine learning algorithms operate fairly. To get a sense of when algorithmic fairness is at issue, imagine a data scientist is provided with data about past instances of some phenomenon: successful employees, inmates who when released from prison go on to reoffend, loan recipients who repay their loans, people who click on an advertisement, etc. and is tasked with developing an algorithm that will predict other instances of these phenomena. While an algorithm can be successful or unsuccessful at its task to varying degrees, it is unclear what makes such an algorithm fair or unfair.
To some, algorithmic fairness is a comparative concept. To see if an algorithm is fair, one looks at how one person or group is assessed by the algorithm as compared to how another person or group is assessed. While such comparative accounts are all committed to the principle that like cases should be treated alike, they differ with regard to what dimensions of likeness they assert matter morally, and why. To others, algorithmic fairness does not have this comparative dimension. Instead, an algorithm is unfair if it is inaccurate or if its processes are obscure, for example. Additional fairness issues focus specifically on the data from which an algorithm is developed. One might wonder whether the only obligations regarding the use of data are epistemic in nature, requiring accurate and representative data, or if, instead, data scientists should be concerned about the ways in which reliance of accurate data perpetuate unfairness.
In what follows, each of these possible fairness issues, and others, are addressed. Section 1 contains an introduction to the topic, and to an important real-world controversy that has been a focus of much scholarship in this area. Section 2 discusses comparative accounts of algorithmic fairness. Section 3 then turns to non-comparative accounts of algorithmic fairness. Section 4 focuses on moral and epistemic issues relating to data collection and use. Finally, Section 5 turns to a specific conceptual problem. Because machine learning algorithms will identify traits that are correlated with legally protected traits like race and sex, scholars wonder when and why these correlated traits should be considered as “proxies” for the protected traits.