- edited by
- Stella Wernicke
Reviewers: Two anonymous reviewers
Editorial Assistant: Maren Giersiepen
When AI determines medical treatments or prison sentences, we may find comfort in the idea of a “human touch” in decisions, for instance, in the form of human oversight mandated in AI regulation. But does human oversight add something meaningful or does it merely serve to convince us to trust AI?
Imagine there are two modern hospitals in town. Because of their advanced technology – and a serious staffing shortage – both rely on the AI tool “DocBot” to develop treatment plans for patients. The first hospital, “AutoMedica,” fully automates treatment plans using DocBot. The second, “BetterTogether,” has doctors review and possibly adapt the treatment plans proposed by DocBot. Which hospital would you trust more, and why?
The promises and risks of AI
AI-based systems are increasingly shaping our lives. AI promises to improve efficiency and accuracy in fields like healthcare, where it can speed up diagnoses and help address shortages in psychotherapy counseling by offering chatbot-based consultations. AI may also support judicial systems by automating routine tasks and expediting case processing [3].
However, these promises come with risks. AI-based systems are prone to errors, as users of generative AI tools like ChatGPT will have noticed. AI can also produce erroneous or ethically problematic outputs at a scale and speed beyond human capacity. Worse, AI decision-making processes are often opaque, making it difficult to understand why a system produces flawed or biased results [4]. Concerns about such risks may consequently also undermine public trust in AI. Even if DocBot can help individuals receive faster treatment in the hospital, the risk of potentially inaccurate and unexplainable outputs may make individuals distrust its use.
Human oversight in AI regulation
Addressing risks in AI use is a key objective of AI regulation, such as the EU AI Act, which came into force in August 2024 and will be incrementally implemented over the coming years. To mitigate risks, the AI Act establishes a set of requirements for those who develop AI systems or have them developed and market them under their name, and those who deploy and operate AI systems. The AI Act classifies AI systems according to their level of risk, and most of its requirements fall on “high-risk” AI. What counts as high-risk is listed in the Annexes of the AI Act and currently includes contexts such as hiring, law enforcement, or health care. For high-risk systems and as a means to mitigate AI risks, the AI Act mandates, for example, a risk evaluation before they are placed on the market, transparency and documentation requirements, as well as monitoring throughout their life cycle. Additionally, Article 14 of the AI Act requires effective human oversight of high-risk AI systems.
Human oversight personnel will be responsible for detecting errors, discriminatory outcomes, and other AI outputs that pose a risk to the health, safety, or fundamental rights of humans. This risk mitigation via human oversight should also contribute to trust-building in AI and highlights a strategic objective of the AI Act: to implement factually trustworthy AI systems, encouraging individuals and organizations to adopt AI (see also Recital 1 of the AI Act, which highlights the overarching goal of the AI Act of “promoting the uptake of human-centric and trustworthy artificial intelligence”). This, in turn, should help unlock AI’s economic and social potential [5]. To be compliant with the AI Act, hospitals using DocBot would thus need to implement human oversight with the goal of humans mitigating risks, such as by detecting and correcting flawed outputs of DocBot. This risk mitigation should then also contribute to building trust in the use of the AI for patients, hospital management, and the broader public.
Does human oversight mitigate AI risks?
While there are promising examples of human oversight mitigating risks of AI use [11], its effectiveness remains uncertain [9], [12]. Research indicates that people have difficulties distinguishing between accurate and inaccurate outputs, especially as AI systems become more advanced [12], [13]. Users of generative AI tools may recognize this challenge – it is becoming increasingly difficult to determine whether often persuasive AI-generated outputs are in fact correct. Even more concerning, human oversight can sometimes introduce errors into AI outputs that were accurate or fair [14]. Overall, humans tend to both over-rely and under-rely on AI [2], [9] – the AI Act even explicitly mentions the risk of humans over-relying on AI [15].
Consequently, human-AI collaboration rarely leads to outcomes that are better than the AI operating alone [12], [16]. In tasks where AI systems exceed human expertise, joint decisions may lead to slight improvements over human-only decisions, but overall, human involvement tends to reduce the accuracy of AI-based decisions [12], [13]. This may not be surprising: when the use of AI systems is considered, it is because they have the potential to outperform humans. If that is the case, how are humans supposed to reliably determine when they should override or ignore AI outputs [9]? If DocBot accurately diagnoses patient conditions in 90% of cases and human doctors in 80% of cases, how often do we believe doctors will be able to tell apart right from wrong DocBot outcomes?
Does human oversight build trust in AI?
In a world where AI-based systems increasingly influence decisions about medical treatments and even sentencing in criminal justice, we may value a “human touch” in these processes. It is great if the human doctor overseeing DocBot reduces the risks of using the AI system, but human involvement may also simply make us feel better about using the AI.
In fact, while evidence on the effectiveness of human oversight in mitigating AI risks remains limited, there is strong evidence that adding a human to AI operations can increase the willingness to adopt AI. A growing number of studies indicate that individuals affected by AI-decisions and the broader public who imagine that AI could soon be used in high-risk domains feel more comfortable with AI-based systems when a human is involved in the process [7], [8].
There are many reasons why adding a human can lead to more positive reactions regarding the use of AI in various use cases [17]. The underlying assumption is that human involvement contributes something valuable, making human-AI collaboration seem superior to AI operating alone, which can foster trust in the use of AI [18]. In line with the AI Act’s risk-mitigation goal, individuals may be more inclined to trust the use of AI if they believe that human oversight can prevent problematic outcomes. Additionally, in certain tasks – especially those involving interpersonal interactions or subjective judgment – people tend to assume that humans possess superior skills [7]. Research also suggests that humans are perceived to be better at considering individual circumstances and demonstrating empathy [19]. Moreover, some studies indicate that people believe it is easier to communicate with a human and potentially persuade them to change a decision [20]. Furthermore, the presence of a human can make AI-driven processes feel more transparent – if DocBot presents a diagnosis, patients may assume they can request an explanation from the doctor [17]. Similarly, having a human involved may suggest that someone can be held accountable if something goes wrong [17]. Even when the human role is purely symbolic, their mere presence can make AI decision-making feel more familiar – much like displaying an image of a credit card when paying with a smartphone [21].
Clearly, many of these reasons reflect attitudes towards AI and humans that may be subject to change over time. If there is one consistent trend in the technologies underlying AI tools, it is that they come with increasing capabilities, such as AI chatbots that generate ever more human-like interactions. With every new update and every new user, familiarity with AI will grow, and this will likely affect human attitudes towards AI and towards having humans involved in AI operations.
The risk of meaningless human oversight
While human oversight may (currently) increase the willingness to adopt AI, it may still not be meaningful. Situations where human oversight increases AI adoption without substantially changing anything for the better can have harmful consequences. For example, DocBot diagnoses may not become more accurate simply because a doctor approves them, yet patients may place greater trust in the diagnoses. Similarly, DocBot may produce results that are unexplainable even to the doctor, but patients may still believe that the doctor can explain diagnoses and treatment plans, thus leading to more patient trust in the use of DocBot. Likewise, if the doctor has no authority to change DocBot’s decisions, the expectation that having the doctor oversee AI enhances the contestability of AI-based decisions is overstated.
As a result, even if human oversight does not actually make anything substantially better than AI operating alone, the mere presence of a human can increase the perceived trustworthiness of AI, foster unwarranted trust, and contribute to its wider adoption [6].
Towards a broader view on effective human oversight
Human involvement in AI operations needs to have a meaningful impact. To achieve this, we must refine our approach to designing human oversight, which requires that we shape regulations and human oversight practice with interdisciplinary perspectives – including expertise regarding the human factors in human-AI interaction [22]. As AI-based systems become embedded in society, regulating and designing these technologies requires acknowledging that there are strong interdependencies between technology, people, and the context in which they interact [10]. Yet, discussions on AI governance often focus primarily on technological optimization, regulatory compliance, and high-level ethical debates. Thus, computer scientists, legal experts, and ethicists dominate these conversations, sometimes complemented by professionals from specific fields of application (e.g., doctors in medical AI). Despite widespread claims about “human-centered” AI design, experts in human factors – such as psychologists and human-computer interaction researchers – are rarely integral to these discussions [17].
The result? Regulations such as the AI Act mandate human oversight but do so in a way that is, at best, empirically naïve. For example, the law focuses almost exclusively on risk mitigation, for which there is limited support of its effectiveness, while overlooking the broader roles human oversight can play and that may foster trust in AI use [17]. This narrow approach risks implementing human oversight that fails not only to reduce AI-related risks but also fails to realize other benefits of human involvement (e.g., transparency, contestability). Had expertise on the human factors in human-computer interaction been integrated earlier in drafting the AI Act, the law could have provided clearer guidance on what human oversight should achieve and how to ensure its effectiveness [10], [15], [17].
While amending Article 14 of the AI Act may be unrealistic, it is possible to ensure that designing human oversight of AI in practice integrates human factors expertise as early as possible. This requires two key steps:
- Clarifying the purpose of human oversight, defining human responsibilities, embedding human oversight in an overall risk-management regime: This requires answering basic questions such as: What is the role of human oversight in a given context? Should oversight personnel focus on detecting errors and biases, enhancing transparency, or ensuring contestability? It is also important to acknowledge that human oversight is only one (albeit potentially the final) layer of risk mitigation in the use of AI. Human oversight must be embedded within a broader risk-management framework, with the understanding that it will not be possible to technically mitigate all risks that AI will introduce – particularly not through human oversight personnel who are expected to serve as the last line of defense against every remaining potential AI risk [23].
- Designing the conditions for effective human oversight: Once the purpose of human oversight is established, we must determine how to design AI systems to support it, how to train and prepare oversight personnel, and how to create suitable working conditions to ensure oversight is meaningful and effective [10]. Optimally, all of this needs to align with the current state of research in the literature on the human factors of interacting with AI systems [15].
Human oversight of AI beyond the “human touch”
For human oversight of AI to be meaningful, it must go beyond merely adding a “human touch.” Human oversight must be well-defined and grounded in empirical evidence. Future policy efforts, research, and practical implementation should first clarify the specific role humans are meant to play in overseeing AI. Then, it is possible to optimize human-AI collaboration to fulfill that role – instead of simply creating the illusion of meaningful oversight. This approach can help ensure that human oversight adds value and leads to well-founded trust in AI use, rather than giving us a false sense of safety just because a human is present. Only when the doctor contributes something meaningful to the use of DocBot should we feel more confident in a hospital that has a human overseeing DocBot.
Bibliography
[1] E. C. Stade et al., “Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation,” npj Mental Health Res, vol. 3, no. 1, p. 12, Apr. 2024, doi: 10.1038/s44184-024-00056-z.
[2] S. Gaube et al., “Underreliance harms human-AI collaboration more than overreliance in medical imaging,” Nov. 22, 2024, Open Science Framework. doi: 10.31219/osf.io/4wv8j.
[3] D. Barysė and R. Sarel, “Algorithms in the court: does it matter which part of the judicial decision-making is automated?,” Artif Intell Law, vol. 32, no. 1, pp. 117–146, Mar. 2024, doi: 10.1007/s10506-022-09343-6.
[4] J. Burrell, “How the machine ‘thinks’: Understanding opacity in machine learning algorithms,” Big Data & Society, vol. 3, no. 1, Jan. 2016, doi: 10.1177/2053951715622512.
[5] J. Laux, S. Wachter, and B. Mittelstadt, “Trustworthy artificial intelligence and the European Union AI act: On the conflation of trustworthiness and acceptability of risk,” Regulation & Governance, vol. 18, no. 1, pp. 3–32, Jan. 2024, doi: 10.1111/rego.12512.
[6] N. Schlicker, A. Uhde, K. Baum, M. C. Hirsch, and M. Langer, “How do we assess the trustworthiness of AI: Introducing the Trustworthiness Assessment Model (TrAM),” PsyArXiv, 2023, doi: 10.31234/osf.io/qhwvx.
[7] M. Langer and R. N. Landers, “The future of artificial intelligence at work: A review on effects of decision automation and augmentation on workers targeted by algorithms and third-party observers,” Computers in Human Behavior, vol. 123, p. Article 106878, Oct. 2021, doi: 10.1016/j.chb.2021.106878.
[8] M. Reis, F. Reis, and W. Kunde, “Influence of believed AI involvement on the perception of digital medical advice,” Nat Med, vol. 30, no. 11, pp. 3098–3100, Nov. 2024, doi: 10.1038/s41591-024-03180-7.
[9] B. Green, “The flaws of policies requiring human oversight of government algorithms,” Computer Law & Security Review, vol. 45, p. Article 105681, Jul. 2022, doi: 10.1016/j.clsr.2022.105681.
[10] S. Sterz et al., “On the quest for effectiveness in human oversight: Interdisciplinary perspectives,” in The 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro Brazil: ACM, Jun. 2024, pp. 2495–2507. doi: 10.1145/3630106.3659051.
[11] M. De-Arteaga, R. Fogliato, and A. Chouldechova, “A case for humans-in-the-loop: Decisions in the presence of erroneous algorithmic scores,” in Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, Honolulu HI USA: ACM, Apr. 2020, pp. 1–12. doi: 10.1145/3313831.3376638.
[12] T. Rieger, L. Onnasch, E. Roesler, and D. Manzey, “Why highly reliable decision support systems often lead to suboptimal performance and what we can do about it,” IEEE Trans. Human-Mach. Syst., vol. 55, no. 5, pp. 736–745, Oct. 2025, doi: 10.1109/THMS.2025.3584662.
[13] M. Vaccaro, A. Almaatouq, and T. Malone, “When combinations of humans and AI are useful: A systematic review and meta-analysis,” Nat Hum Behav, vol. 8, no. 12, pp. 2293–2303, Oct. 2024, doi: 10.1038/s41562-024-02024-1.
[14] B. Green and Y. Chen, “Disparate interactions: An algorithm-in-the-loop analysis of fairness in risk assessments,” in Proceedings of the Conference on Fairness, Accountability, and Transparency, Atlanta GA USA: ACM, Jan. 2019, pp. 90–99. doi: 10.1145/3287560.3287563.
[15] J. Laux and H. Ruschemeier, “Automation bias in the AI Act: On the legal implications of attempting to de-bias human oversight of AI,” 2025, arXiv. doi: 10.48550/ARXIV.2502.10036.
[16] P. Hemmer, M. Schemmer, N. Kühl, M. Vössing, and G. Satzger, “Complementarity in human-AI collaboration: concept, sources, and evidence,” European Journal of Information Systems, p. Article 2475962, Aug. 2025, doi: 10.1080/0960085X.2025.2475962.
[17] R. Crootof, M. E. Kaminski, and W. N. Price Ii, “Humans in the loop,” Vanderbilt Law Review, vol. 429, 2023, Accessed: Mar. 07, 2025. [Online]. Available: https://scholarship.law.vanderbilt.edu/vlr/vol76/iss2/2
[18] N. Schlicker, K. Baum, A. Uhde, S. Sterz, M. C. Hirsch, and M. Langer, “How do we assess the trustworthiness of AI? Introducing the trustworthiness assessment model (TrAM),” Computers in Human Behavior, vol. 170, p. 108671, Sep. 2025, doi: 10.1016/j.chb.2025.108671.
[19] C. Longoni, A. Bonezzi, and C. K. Morewedge, “Resistance to medical artificial intelligence,” Journal of Consumer Research, vol. 46, no. 4, pp. 629–650, 2019, doi: 10.1093/jcr/ucz013.
[20] H. Lyons, T. Miller, and E. Velloso, “Algorithmic decisions, desire for control, and the preference for human review over algorithmic review,” in 2023 ACM Conference on Fairness, Accountability, and Transparency, Chicago IL USA: ACM, Jun. 2023, pp. 764–774. doi: 10.1145/3593013.3594041.
[21] K. Brennan-Marquez, K. Levy, and D. Susser, “Strange loops: Apparent versus actual human involvement in automated decision making,” Berkeley Technology Law Journal, vol. 34, no. 3, pp. 745–772, 2019.
[22] S. Gaube et al., “Keeping an eye on AI: A framework for effective human oversight of AI systems,” 2026, arXiv. doi: 10.48550/ARXIV.2605.16278.
[23] J. Laux, “Institutionalised distrust and human oversight of artificial intelligence: towards a democratic design of AI governance under the European Union AI Act,” AI & Soc, vol. 39, no. 6, pp. 2853–2866, Dec. 2024, doi: 10.1007/s00146-023-01777-z.
Image sources
Image 1: Photo by ThisisEngineering on Unsplash
Image 2: Photo by janilson furtado on Unsplash
Image 3: Photo by Vitaly Gariev on Unsplash
Funding
Markus Langer received funding from the German Research Foundation (DFG) via TRR 248 – Foundations of Perspicuous Software Systems (CPEC, Project 389792660) and by the Daimler and Benz Foundation (project TITAN - Technologische Intelligenz zur Transformation, Automatisierung und Nutzerorientierung des Justizsystems, grant no. 45-06/24).
Kevin Baum received funding from the European Regional Development Fund (ERDF) and Saarland for Towards a Center for European Research in Trusted Artificial Intelligence (ToCERTAIN), from the German Research Foundation (DFG) via TRR 248 – Foundations of Perspicuous Software Systems (CPEC, Project 389792660), and from the German Federal Ministry of Education and Research (BMBF) for Multi-Level Abstractions and Causal Modeling for Enhanced Reinforcement Learning (MAC-MERLin, Grant 01IW24007).
Johann Laux received funding from a British Academy Postdoctoral Fellowship (grant no. PF22\220076) via the Emerging Laws of Oversight project.
- Log in to post comments