“It's about making good use of modern methods while also understanding their limitations.”

Digital methods open up new opportunities for research and teaching in the humanities and cultural studies—but they also present researchers with new challenges when it comes to handling data. In the Digital Humanities RUHR project, statistician Dr. Henrike Weinert demonstrates why data literacy, basic statistical skills, and the question of responsibility in algorithmic systems are central to this field.
You’re a statistician, but you’re involved in the “Digital Humanities RUHR” initiative. How do those two fit together, and what is the initiative actually about?
Digital methods have become indispensable in the humanities and cultural studies as well—for example, for analyzing large amounts of data. The three UA Ruhr universities have therefore joined forces in the “Digital Humanities RUHR” project and developed a joint certificate program to provide students—even those in fields less focused on data—with key competencies for academic work and their future careers. The Foundation for Innovation in Higher Education has funded the project—which I coordinate on behalf of TU Dortmund University—for just over two years through the end of June, and we were already able to award the first certificates at the beginning of the year. To earn the certificate, students must complete introductory and advanced modules and present their own projects in a final colloquium.
Each university contributes a specific focus to the project; for us at TU Dortmund, it is “Algorithmic Accountability.” Among other things, this addresses the question of who actually bears responsibility for decisions that have been influenced by the use of algorithms. The interdisciplinary seminar we’ve developed on this topic involves the fields of statistics, data science, journalism, and computer science. In this seminar, students work in small interdisciplinary groups to develop journalistic products—such as an article—while exploring topics like the evaluation of statistical models or how algorithms function. In addition, as part of “Digital Humanities RUHR,” we’ve developed new learning materials and tested new teaching formats.
What needs do you currently see in the areas of data literacy and research data management?
Perhaps I say this because I’m a statistician, but for me, basic statistical skills are essential. Everything hinges on data quality—and that applies just as much to AI applications. Proper data collection is crucial. At the same time, one must transparently explain when data could not be collected under ideal conditions. It is possible to work with such data, but one must clearly identify the existing uncertainties and make it clear when results are more like indications than reliable findings. Basic knowledge is often still lacking here: How good does data need to be to make certain statements based on it? What conclusions are even possible, depending on the study design? Many of these questions actually pertain to statistics, but are often not associated with it. That’s why we found the term “data literacy” helpful—it encompasses much of this without the often-negative perception associated with statistics.
A common misconception is the idea that one simply needs enough data and the rest will follow naturally. In reality, data always provides only an incomplete picture of reality, and uncertainties are always part of the picture. These must be clearly communicated. The phrase “The data speaks for itself” is never true in practice. Another important point is the entire process surrounding data: from formulating the research question through operationalization and data collection to documentation and evaluation. You have to clarify why the data was originally collected, what its quality is, and whether it’s even suitable for your specific research question. Only then does the actual analysis begin. Data literacy therefore also means understanding fundamental concepts—such as when an average makes sense and when, for example, a median is a better fit. Especially with more complex models or AI systems, it becomes even harder to understand how results are derived. That is why, above all, we need greater mutual understanding: Not everyone has to become a data scientist, but experts must be able to discuss data with one another. This also includes critical thinking and ethical questions—that is, what one can, may, and should do with data. In some cases, it can even be problematic not to use data or AI, such as when they demonstrably enable better analyses of X-ray images.
What potential do you see in greater digitization for research in the humanities and cultural studies?
The term “digital humanities”—that is, the application of digital technologies and methods in the humanities and cultural studies—is certainly controversial. Some say that good research in the humanities is becoming increasingly digital anyway, so there’s no need to emphasize it separately. Others see it as something special precisely because many of these disciplines have traditionally relied little on data, data collection, or statistical methods. Digital methods such as text mining open up new possibilities here because computers can analyze large amounts of text much more quickly. I see great potential in this, but at the same time, many unanswered questions. Precisely because digital tools enable us to work with significantly larger amounts of data, we must pay closer attention to issues of proper application, uncertainty, and interpretation. Our project therefore focused precisely on this aspect of teaching: How can we help students in the humanities—who generally have little exposure to statistics—develop basic data literacy? At the University of Duisburg-Essen, for example, an introductory course was restructured so that students can no longer pursue certain degree programs without having worked with datasets and digital methods at least once.
At the same time, it has become apparent that there is a lack of even very basic computer skills in some cases. At Ruhr University Bochum, basic courses have therefore been developed to teach simple digital skills. In the long term, the goal is to be able to use modern methods effectively while also understanding their limitations—especially when it comes to AI systems, which often still seem like a black box. This also includes ethical questions that remain unresolved in many cases. Particularly with regard to journalistic reporting, improved data literacy can also have beneficial societal impacts in the long term—for example, by enabling citizens to evaluate information more critically and not to believe every claim on social media without scrutiny.
About the Person:
- Since 2004, research assistant and instructor for special assignments in the Departments of Statistics and Mathematics
- Since 2020, research assistant at the Center for Data Science and Simulation (DoDaS) at TU Dortmund University; director of the “Data Competence Network” data and AI literacy program
- Since 2024, she has been involved in the UA Ruhr initiative “Digital Humanities,” focusing on algorithmic accountability in Dortmund
Dr. Henrike Weinert is featured as a Data Champion because she teaches data literacy even in non-data-intensive degree programs and highlights the added value of research data management for research in the humanities and cultural studies.
Further information:
