Measuring Gender Bias in German Language Generation
Abstract: Most existing methods to measure social bias in natural language generation are specified for English language models. In this work, we developed a German regard classifier based on a newly crowd-sourced dataset. Our model meets the test set accuracy of the original English version. With the classifier, we measured binary gender bias in two large language models. The results indicate a positive bias toward female subjects for a German version of GPT-2 and similar tendencies for GPT-3. Yet, upon qualitative analysis, we found that positive regard partly corresponds to sexist stereotypes. Our findings suggest that the regard classifier should not be used as a single measure but, instead, combined with more qualitative analyses.
Show BibTeX
@inproceedings{DBLP:conf/gi/KraftZFSBU22,
author = {Angelie Kraft and
Hans{-}Peter Zorn and
Pascal Fecht and
Judith Simon and
Chris Biemann and
Ricardo Usbeck},
editor = {Daniel Demmler and
Daniel Krupka and
Hannes Federrath},
title = {Measuring Gender Bias in German Language Generation},
booktitle = {52. Jahrestagung der Gesellschaft f{\"{u}}r Informatik, {INFORMATIK}
2022, Informatik in den Naturwissenschaften, 26. - 30. September 2022,
Hamburg},
series = {{LNI}},
volume = {{P-326}},
pages = {1257--1274},
publisher = {Gesellschaft f{\"{u}}r Informatik, Bonn},
year = {2022},
url = {https://doi.org/10.18420/inf2022\_108},
doi = {10.18420/INF2022\_108},
timestamp = {Mon, 03 Mar 2025 21:05:10 +0100},
biburl = {https://dblp.org/rec/conf/gi/KraftZFSBU22.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}