- edited by
- Stella Wernicke,
- Magdalena Wischnewski
Reviewers: Dr. Marinus Ferreira and one anonymous reviewer
Editorial Assistant: Zoey Chapman
With the growing popularity and increasing capabilities of generative AI, more and more attention is also drawn to AI model biases, in particular, gender biases. But what exactly does bias mean in this context, how should it be addressed, and what role do humans play in all this?
Up until now, countless parents all over the world have had to strain their brains every evening to come up with new creative bedtime stories for their children. Every morning, hundreds of journalists are tasked with putting simple information, such as the results of yesterday’s football game, into words. And every day, countless foreign language teachers try to get shy teenagers to talk so they can have a bumpy conversation in the language they are learning. All this could soon be a thing of the past – thanks to technological advances in the field of generative artificial intelligence (GenAI). GenAI can be defined as “computational techniques that are capable of generating seemingly new, meaningful content such as text, images, or audio from training data” [1]. As such, GenAI is nowadays capable of creating original, individualized stories about each child’s personal hero or heroine and even reading them out loud, including mimicking different voices or leaving dramatic pauses for effect. GenAI is also already frequently used in journalism to automatically generate elaborate articles from data such as sports statistics or interview recordings. Lastly, GenAI may also revolutionize language learning through communication with chatbots especially programmed to improve grammar, expand vocabulary, or correct mistakes, all in the course of a lively, entertaining conversation.
These are just some examples that show how GenAI could be of use in virtually every domain of everyday life in the future. While the novel technology thus holds great promise, it does not come entirely without risks. For instance, chatbots, like the most popular model ChatGPT, are prone to “hallucinate,” that is, completely make up entirely false information – like falsely telling a man who asked ChatGPT for information about his person that he had killed his children [2]. In addition, GenAI models may also be used to strategically spread disinformation, for example, in the context of current armed conflicts like the ones between Russia and Ukraine [3] or the one between Israel and Palestine/Hamas [4]. Another risk, the one that is also at the heart of this article, is the perpetuation of biases and stereotypes in the various forms of content generated by GenAI.
Gender bias in GenAI – how much is too much?
Gender is the most salient identity dimension – it is the one that children distinguish before any other [5], and it is also generally perceived as the most natural category of differentiation [6]. In this light, it is not surprising that gender bias is also one of the, if not the type of bias that is most evident when it comes to GenAI. For instance, text-generating models more often associate emotional or social themes (e.g., ‘fear’ or ‘community building’) with female protagonists, while narratives focused on intellect or competition (e.g., ‘logic’ or ‘fighting’) are mainly linked to male characters [7]. Similar tendencies can also be found in the context of occupations. When asked to write a short story about a pilot and their cabin crew, for example, chatbots will most likely tell the tale of a male pilot and a female flight attendant. Similarly, in a story about a CEO and their employee, chatbots predominantly assign the part of the CEO to a man, while the role of the secretary is attributed to a woman [8]. Even in a fictional story, it seems that a chatbot cannot imagine a woman flying a plane or running a company.
A very similar pattern also emerges when prompting image-generating AI models to create pictures of different occupations. Carpenter, sawing machine operator, construction worker – name any occupation that requires physical strength, and you can be almost certain to get an image of a man. When it comes to nursing, care, and social jobs – you already guessed it – the AI-generated picture will show a woman [9]. One might argue now that this imbalance comes as no surprise. After all, the census shows that carpentry is a largely male-dominated occupation, while nursing is a career most often pursued by women [10]. Yet, text- and image-generating AI models have been shown not only to reproduce but amplify real-life gender inequalities. Male-dominated (and high-prestige) professions, such as management or software development, are predominantly portrayed with disproportionately more men than their real-world gender composition would suggest [11][12]. On the other hand, for female-dominated professions, gender disparities are in some cases amplified, in others reduced, contributing, however, to an overall underrepresentation of women in occupational depictions [8].
The balancing act of addressing bias in AI
Addressing this issue, however, is challenging for a number of reasons. First, “bias” seems to have become a popular buzzword that is tossed around without giving much thought to what it actually refers to. Some employ the term bias for any gender distribution that is not entirely equal. In this sense, however, it is virtually impossible for any AI model not to be biased, as it would require a perfect fifty-fifty share of men and women across all occupations. Others, therefore, refer to AI-generated content as biased only when the gender distributions significantly deviate from real-world gender distributions. If the issue is already complicated regarding tangible facets of gender bias, such as occupational stereotypes, things get even more tricky when it comes to more subtle aspects, such as character traits. Without a census to compare to, it is hard to assess if AI models ‘just’ reflect societal gender roles or if they aggravate potentially harmful stereotypes. In addition, once the focus shifts beyond gender to other demographic characteristics, the problem becomes even less clear-cut. While a fifty-fifty split may at least sound reasonable in the case of gender, there is no equally clear benchmark for broader diversity categories. Some AI systems appear to address this by distributing outputs evenly across a few broad racial groups, but such artificial parity can be arbitrary and may not reflect any real population. At the same time, demographic benchmarks are not unproblematic either, as many AI models are trained on heavily U.S.-centered data and may therefore incorrectly assume American demographic patterns as a global default.
Next to that, the question of how to deal with bias in GenAI is not just one of definition but also a moral one: Should AI models simply represent our social world as is, without glossing over societal inequalities, whether it be gender or other types of biases? Or should GenAI rather take an active role as a driver of social change? In other terms, coming back to the example of occupational biases, should AI models such as ChatGPT keep on imagining a world where doctors, professors, and CEOs are predominantly White, heterosexual men? Or should the novel technology paint a different picture that encourages its users to face social injustices and not be defined by traditional stereotypes?
The companies behind most of the currently popular AI models, such as OpenAI, the creator of ChatGPT, seem to opt for the latter [13]. For instance, models will refuse to carry out certain requests, such as creating images of or stories about ‘stereotypical’ men or women. Yet, trying to tackle societal biases in AI models represents a tightrope act between boosting diversity without overshooting. Google’s Gemini, for instance, provided a vivid example of what can happen when the limit is pushed too far, and the goal of promoting diverse outputs can actually result in (historically) inaccurate information – creating images of female popes, Black Vikings, or Asian World War I soldiers [14]. Attempts to counter gender bias in AI can also lead to asymmetrical overcorrections, for example, in that women are more readily cast in stereotypically masculine roles than the other way around, or in that ChatGPT-4 is more likely to condone violence by a woman against a man than by a man against a woman [15].
When it comes to the potential implications of such biased depictions – whether stereotypical or counter-stereotypical – we know much more about what AI models produce than about what these portrayals actually do to their users. Still, the available evidence suggests that biased outputs can undermine users’ trust in the AI models [16] and may reinforce harmful stereotypes [17]. At the same time, counter-stereotypical portrayals may also hold transformative potential, potentially reducing existing gender stereotypes held by users [18].
The role of humans in AI-generated bias
In most discussions about gender bias in GenAI, AI models are talked about as if they were independent entities, sometimes even as if they were independent persons – for example, referring to ChatGPT as “he” is a slip of the tongue that happens to me almost more often than not. However, as exactly this slip of the tongue illustrates, blaming chatbots for reproducing some degree of gender bias seems hardly fair when the gender stereotypes ingrained in our brains still lead us to assume that a clever and knowledgeable chatbot must be male (especially in contrast to the helpful, service-oriented voice assistants Siri and Alexa). Further, only talking about the bias in the AI models makes it easy to overlook the bias that is at the root – namely, the bias in the training data. AI models are fed with tons and tons of all kinds of digital information available on the internet, most of which is far from neutral or balanced. An exploration of the search engine Bing, for instance, has shown that search queries including character traits such as sensitivity or emotionality more often yield pictures of women, while attributes like ambitiousness or intelligence produce images of men [19]. Next to the data the models are trained on, there is also a bias in who is training these models, with women working in AI still being outnumbered by twice as many men [20].
Lastly, tackling gender bias in GenAI can only really be fruitful if the gender bias held by its users is also addressed. For instance, in the context of career recommendation AI models, developing a balanced model that does not base its suggestions on the gender of the inquirer has been achieved without much of an issue. Yet, when testing the model with real users, it was discovered that users actually prefer the biased model. Their own inherent notions of what roles they should fulfill based on their gender lead them to give counter-stereotypical suggestions a significantly worse rating than those fitting the gender stereotypes [21]. Virtually no aspect of our social world is entirely free of gender bias – what we perceive as “normal” is biased at heart. For this reason, it is also extremely difficult for users to spot biases in GenAI, even when they are first alerted to the biased nature of the AI-generated content [22]. Simply blaming the biased bot, thus, definitely falls short of acknowledging the complex interdependencies of AI models and human gender biases – be it in the training data, the model design choices, or in the stereotypes held by developers. Addressing gender bias in GenAI, therefore, requires action at multiple levels, ultimately including users themselves. Rather than taking AI outputs at face value, especially when they concern social roles, occupations, or personality traits, users should ask who is being portrayed, how they are being portrayed, and who might be missing. Used critically in this way, GenAI can become not only a mirror of existing biases but also a tool for questioning them.
Bibliography
[1] S. Feuerriegel, J. Hartmann, C. Janiesch, and P. Zschech, “Generative AI,” Bus. Inf. Syst. Eng., vol. 66, no. 1, pp. 111–126, Feb. 2024, doi: 10.1007/s12599-023-00834-7.
[2] I. Rahman-Jones, “ChatGPT falsely told man he killed his children,” BBC, Mar. 21, 2025. Accessed: Mar. 26, 2025. [Online]. Available: https://www.bbc.com/news/articles/c0kgydkr516o
[3] J. Twomey, D. Ching, M. P. Aylett, M. Quayle, C. Linehan, and G. Murphy, “Do deepfake videos undermine our epistemic trust? A thematic analysis of tweets that discuss deepfakes in the Russian invasion of Ukraine,” PLOS ONE, vol. 18, no. 10, p. e0291668, Oct. 2023, doi: 10.1371/journal.pone.0291668.
[4] G. C. Stănescu, “Fake News in Times of Conflict: AI-Driven Disinformation during the Israel-Iran Crisis,” Nov. 2025, doi: 10.5281/ZENODO.17603896.
[5] D. M. Mackie, D. L. Hamilton, J. Susskind, and F. Rosselli, “Social psychological foundations of stereotype formations.,” in Stereotypes and Stereotyping, C. N. Macrae, C. Stangor, and M. Hewstone, Eds., Guilford Press, 1996, pp. 41–78.
[6] N. Haslam, L. Rothschild, and D. Ernst, “Essentialist beliefs about social categories,” Br. J. Soc. Psychol., vol. 39, no. 1, pp. 113–127, 2000, doi: 10.1348/014466600164363.
[7] I. C. E. Van Blerck, E. S. De Lima, M. M. E. Neggers, and T. Calders, “Unveiling gender bias in LLM-generated hero and heroine narratives,” Entertain. Comput., vol. 55, p. 100972, Sep. 2025, doi: 10.1016/j.entcom.2025.100972.
[8] D. A. Voutyrakou, G. Katsiampoura, and C. Skordoulis, “Are AI tools biased storytellers? Examining gender-bias in AI-generated narratives,” Edelweiss Appl. Sci. Technol., vol. 9, no. 4, pp. 1032–1049, Apr. 2025, doi: 10.55214/25768484.v9i4.6169.
[9] L. Girrbach, S. Alaniz, G. Smith, and Z. Akata, “A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models,” Sep. 29, 2025, arXiv: arXiv:2503.23398. doi: 10.48550/arXiv.2503.23398.
[10] M. Davern, R. Bautista, J. Freese, P. Herd, and S. L. Morgan, “General Social Survey 1972-2022.” NORC, 2024. [Online]. Available: https://gss.norc.org/us/en/gss/get-the-data.html
[11] O. Menis–Mastromichalakis, G. Filandrianos, M. Symeonaki, G. Stamatopoulou, D. Parsanoglou, and G. Stamou, “Gender bias in machine learning: insights from official labour statistics and textual analysis,” Qual. Quant., vol. 60, no. 1, pp. 619–653, Feb. 2026, doi: 10.1007/s11135-025-02261-0.
[12] F. Bianchi et al., “Easily accessible text-to-image generation amplifies demographic stereotypes at large scale,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, in FAccT ’23. New York, NY, USA: Association for Computing Machinery, Jun. 2023, pp. 1493–1504. doi: 10.1145/3593013.3594095.
[13] OpenAI, “Evaluating fairness in ChatGPT.” Accessed: Mar. 29, 2025. [Online]. Available: https://openai.com/index/evaluating-fairness-in-chatgpt/
[14] G. D. Vynck and N. Tiku, “Google takes down Gemini AI image generator. Here’s what you need to know.,” The Washington Post, Feb. 23, 2024. Accessed: Mar. 29, 2025. [Online]. Available: https://www.washingtonpost.com/technology/2024/02/22/google-gemini-ai-i…
[15] R. A. Fulgu and V. Capraro, “Surprising gender biases in GPT,” Comput. Hum. Behav. Rep., vol. 16, p. 100533, Dec. 2024, doi: 10.1016/j.chbr.2024.100533.
[16] C. Chen and S. S. Sundar, “Is this AI trained on credible data? The effects of labeling quality and performance bias on user trust,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, in CHI ’23. New York, NY, USA: Association for Computing Machinery, Apr. 2023, pp. 1–11. doi: 10.1145/3544548.3580805.
[17] A. Gaba, E. Wall, T. R. Babu, Y. Brun, K. Hall, and C. X. Bearfield, “Bias, accuracy, and trust: Gender-diverse perspectives on large language models,” Jul. 08, 2025, arXiv: arXiv:2506.21898. doi: 10.48550/arXiv.2506.21898.
[18] M. Sim, N. G. Brigham, T. Kohno, T. E. S. Charlesworth, and A. Caliskan, “Biased AI outputs can impact humans’ implicit bias: A case study of the impact of gender-biased text-to-image generators,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Oct. 2025, pp. 2375–2386. doi: 10.1609/aies.v8i3.36723.
[19] J. Otterbacher, J. Bates, and P. Clough, “Competent men and warm women: Gender stereotypes and backlash in image search results,” in Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, Denver Colorado USA: ACM, May 2017, pp. 6620–6631. doi: 10.1145/3025453.3025727.
[20] World Economic Forum, “Global Gender Gap Report 2023,” 2023. Accessed: Nov. 27, 2023. [Online]. Available: https://www3.weforum.org/docs/WEF_GGGR_2023.pdf
[21] C. Wang et al., “When Biased Humans Meet Debiased AI: A Case Study in College Major Recommendation,” ACM Trans. Interact. Intell. Syst., vol. 13, no. 3, pp. 1–28, Sep. 2023, doi: 10.1145/3611313.
[22] S. Williams-Ceci, M. Jakesch, A. Bhat, K. Kadoma, L. Zalmanson, and M. Naaman, “Biased AI writing assistants shift users’ attitudes on societal issues,” Sci. Adv., vol. 12, no. 11, p. eadw5578, Mar. 2026, doi: 10.1126/sciadv.adw5578.
Image sources
Image 1: https://pixabay.com/de/illustrations/fantastisch-android-roboter-cyborg…
Image 2: Generated by the author using DALL·E 3.
Image 3: Screenshot by the author of a ChatGPT-4o conversation.
Image 4: https://www.pexels.com/de-de/foto/mann-hande-laptop-internet-16094041/
- Log in to post comments