Critical Technical Practice for Responsible Medical Image Analysis
Critical Technical Practice for Responsible Medical Image Analysis
By Weina Jin (weinaj@sfu.ca)
2026-07-13
TL;DR: Critical technical practice can safeguard our good intentions in medical image analysis against distortion and manipulation from unscientific and unethical assumptions and values, analogous to strengthening our technical immune systems.
1-min promotional video for this tutorial, recorded for MICCAI Educational Challenge 2026.
Abstract: In this article, I motivate the need for critical technical practice using my research experience in conducting AI technical research on explainable AI for medical image analysis. I summarize the different levels of approaches to conduct critical technical practice, and propose ways forward to collectively conduct critical technical practice to make our techniques more scientifically rigorous, ethical, and responsible.
Table of contents
- Critical Technical Practice for Responsible Medical Image Analysis
Does good intentions of tech for healthcare guarantee good outcomes?
For people working in the medical image analysis field, most of us are motivated by good intentions of using technologies to do good for health. I also have the same motivation when coming into the medical image analysis field: if I can develop techniques that help healthcare professionals and patients, that would be fantastic! How can it possibly go wrong?
Well, it turns out that my initial assumption was wrong. Good intentions don’t guarantee good outcomes. “The road to hell is paved with good intentions,” which proverb I learned the hard way from my PhD research experience. The logic is simple: just as scientific outcomes are hard to achieve and can easily be distorted during the process, the same holds true for good outcomes. I was wrong in naively assuming that no distortion was the default; instead, distortion and manipulation of our well-meaning goals should be assumed to be the default to keep us epistemically vigilant. Given this, we need a strong technical immune system to protect our good intentions, mirroring the self-correcting mechanism of science. I call this immune system critical technical practice, which is to critically examine assumptions and justifications in common technical practices, and understand weaknesses, flaws, and boundaries of a technique.
Why good intentions can be distorted in technical development?
I would like to share my research experience of how I realized my good intentions can be distorted in technical development. When I first started my PhD research in explainable AI (XAI) for medical image analysis, I aimed to conduct an evaluation of different XAI algorithms to test their clinical suitability. I surveyed XAI literature and found that one of the popular evaluation metric is plausibility, which measures the reasonableness of AI explanations according to human knowledge on the given task. Because it is a reasonable and popular metric, I used it in my first XAI paper. Later, when I had the opportunity to work with our clinical collaborator, I clearly remember that the doctor mentioned that when he saw a plausible explanation, he would trust the AI system more; and when he saw an implausible explanation, he interpreted it as a good sign to enable him to see through the AI system to understand the flaws in it. The plausible explanation the doctor commented on, however, actually corresponds to an incorrect AI prediction. Just at that moment, I realized that there’s something wrong with the plausible metric in XAI: the use of the plausibility metric to evaluate and optimize XAI algorithms actually encourages the occurrence of misleading explanations, those that explain AI’s wrong predictions using seemingly plausible explanations.
In our final paper of this project, we pointed out the problem with the plausibility metric, and used another criterion of informative plausibility instead, which means the correlation between explanation plausibility and AI prediction quality. After done all these things, I still felt that there’s something weird: If it’s not difficult to identify problems of the plausibility metric (like in my experience, I just showed doctor the AI explanation and found out its problem), then why the entire XAI community hasn’t found its problem and stopped using it as an XAI metric? To the contrary, plausibility is the most commonly used XAI metric according to a systematic review. This is worrisome because anyone who uses XAI algorithms for good intentions such as making AI more transparent and explainable to clinical users, is likely to end up misleading and manipulating users by using the most popular XAI metric of plausibility, just like what I did in my first XAI paper. In this plausibility case, the unethical consequence of misleading users cannot be attributed to any bad actors. Bad outcome comes exactly from ordinary technical practice by following technical conventions and common practices. “A bad system will beat a good person every time.” So the problem is in the social system of the technical community that sets wrong metric as XAI technical standards. Then why it is possible for a scientific community to make such an obvious mistake without self correcting it timely?
For a while, I didn’t have a good answer to this question until I read two books that tell stories about how the big corporations can manipulate and distort scientific agendas can easily be manipulated by the for-profit powerful:
In the book Soda Science: Making the World Safe for Coca-Cola, anthropologist Susan Greenhalgh tells the story of how industry leader Coca-Cola collaborated with academia to conduct real scientific research that advocated exercise, not calorie restraint, as the priority solution for obesity. This distorted research agenda influenced public health policies on obesity and public understanding on diet and lifestyle in favor of the needs and profits of soda industry.
In the book Ghost-Managed Medicine: Big Pharma’s Invisible Hands, Sergio Sismondo detailed how ordinary research activities can be shaped to server the private, not public, interests.
“Pharmaceutical companies sustain large networks to gather, create, control and disseminate information. They provide the pathways that carry this information, and the energy that makes it move. Through bottlenecks and around curves, knowledge is created, given shape by the channels it navigates. Pharma companies create medical knowledge and move it to where it is most useful; much of it is perfectly ordinary knowledge that happens to support their marketing goals. But because of the companies’ resources, their interests and their levels of control, they become key shapers of almost all medical terrains.”
It seems that big food, big pharma, and big tech corporations have the power to shape scientific communities according to their for-profit agendas. But how does the manipulation of scientific and medical agendas so successful? People in the scientific communities are researchers and scientists. Don’t they know they are being manipulated?
There is another paragraph in Ghost-Managed Medicine that detail how the manipulation appears as a natural, inevitable, and scientific process:
“Together, the many elements that pharmaceutical companies shape, adjust and assemble constitute markets. These markets are new creations, but because they draw together medical science and health needs they take on an appearance of necessity. They look like entities that have emerged whole from just below the social surface. The goal of pharma’s assemblage marketing is to establish conditions that make specific diagnoses, prescriptions and purchases as obvious and frequent as possible. Ideally, all of the elements of a market can be directed towards the same issues, claims and facts, so that the drugs sell themselves. Pharma companies can then recede into the background, and apply only minimal pressure when needed.”
In The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence, by foregrounding the underlying ideologies of the AGI agenda, Timnit Gebru and Émile P. Torres describe similar phenomenon of how the powerful parties set the mainstream agenda in the AI community that appears to be “a natural progression”:
‘This investment has succeeded in legitimizing the AGI race such that many students and practitioners who may not be aligned with TESCREAL utopian ideals are working to advance the AGI agenda because it is presented as a natural progression in the field of AI. In the same way that first-wave eugenicists and race scientists sought and achieved academic legitimacy for their research (Saini, 2019), TESCREALists have created a veneer of scientific authority that makes their ideas more palatable to uncritical audiences, and thus have succeeded in influencing research and policy directions in the field of AI. First-wave eugenics proved to be ineffective and catastrophic. But as Jean Gayon and Daniel Jacobi signify with the term “eternal return of eugenics,” eugenic ideals keep on being repackaged in different forms [129]. The AGI race is yet another attempt, diverting resources and attention away from potentially useful research directions, and causing harm in the process of trying to achieve a techno-utopian ideal crafted by self appointed “vanguards” of humanity.’
So the mechanism of how powerful corporations manipulate scientific agenda is more insidious and complicated than I assumed. Improving our technical immune system against the conventional formed dysfunctional and problematic practices means we need to understand the mechanisms at the roots, which is related to how power works in the sociotechnical system of scientific community. I wrote this paper Making Power Explicable in AI: Analyzing, Understanding, and Redirecting Power to Operationalize Ethics in AI Technical Practice with my co-authors to detail such mechanisms, using the XAI field as a case study. In short, as the picture illustrates, if unjust power has penetrated and polluted the roots that makes the whole technical conduct appear to be natural, our technical immune system should also function in all the pathways from roots to leaves to defend against the pollution, by critically reflecting the taken-for-granted assumptions and setting preventative guardrails. Next, I will describe my proposed approach to achieve this goal.

Critical technical practice in medical image analysis: a preliminary approach
Here I would like to describe my current approach to critical technical practice in medical image analysis (MIA). I hope this can provide some concrete idea to inspire you to consider how to strengthen the technical immune system in your task or your own subfield in MIA. Note this is a constantly evolving and dynamic approach. Since strengthening our technical immune system is a systematic and collective effort, a key strategy along the way is to find allies and collaborators to work together toward this goal. This is the motivation for me to write this article.
I categorize approaches that tackle the critical aspects in technical practice into three levels: the high level, middle level, and low level technical practice according to the scope of the problem. This can correspond to the roots, trunk/branches, and leaves in the above tree metaphor.
High level technical practice
Criticality in high level technical practice refers to questioning the overarching goals of technical development, to ask more “why” questions. As technical people, we probably ask more “how” questions than the “why” questions.The “how” questions enable us to solve problems with techniques, i.e, make better hammers, while the “why” questions challenge the fundamental motivation and the overall assumptions in the technical problem formulation, i.e, questioning if we treat everything as nails that can be solved by our hammer, or challenging the motivation for us to build hammer (e.g.: why not building screwdriver or other tools, or why do we need to build tools).
For example, in our paper AI for Just Work: Constructing Diverse Imaginations of AI beyond “Replacing Humans”, we urge the AI community to ask more “why” questions by pointing out limitations in the existing answers to the “why develop AI” question, including outperforming humans, improving efficiency, and freeing human labor from toil. Similarly for medical image analysis tasks, we can question the assumptions that motivate MIA techniques in the first place, such as improving clinical decision quality and efficiency, or reducing medical errors. For example, in “Ethical Medical Image Synthesis”, we point out that the motivation of “overcoming human error” assumes that computational models can be epistemically more capable than healthcare professionals, which is a manifestation of epistemic injustice. This assumption influences the low level technical practice of how the evaluation paradigm of computational models is set: evaluation that emphasizes model strengths but ignores model weaknesses and limitations (v1, page 40-41).
We can also question the fundamental assumptions in the current healthcare systems. For example:
- Is improving efficiency really the optimal way to improve healthcare quality?
- Are there any other non-technical or less technical ways that may be more effective to improve healthcare quantity?
- In her book Ordinary Medicine, Sharon R. Kaufman rethinks medical goals and questions the fundamental assumptions that drive healthcare innovation and services of “more is better”.
- In her book The Logic of Care, Annemarie Mol questions the fundamental concepts of “good care” and “patient choice”.
- In her book The Impossible Clinic, Ariane Hanemaayer provides a critical inspection of evidence-based medicine.
These fundamental questions may be beyond the technical domain and more into the social and philosophical realm, and may not be easily answered. But the process of seeking answers to these big questions itself can be rewarding and intellectually fulfilling. More importantly, it enables us to foreground the underlying big picture, which may be set to appear as natural and inevitable but not necessarily so.
Middle level technical practice
Middle level technical practice refers to the set of assumptions, problem formulations, and problem-solving paradigms that outline common technical approaches in a task or field. For example, in MIA field, regarding the most common research paradigm, it usually involves data-driven model building, with the pipeline of having a medical imaging problem to solve, a deep learning based model to design and train, and evaluation methods to test the model performance in comparison with benchmark or baseline models. Criticality in middle level technical practice means to critically reflect and examine the scientific rigor and responsibility in these practices regarding different MIA tasks such as image recognition, segmentation, and registration.
As an example of criticality for middle level technical practice, in our paper Ethical Medical Image Synthesis, we provide a critical reflection of the technical practice of the medical image synthesis task to make it more ethical and scientifically rigorous. The critical reflection focuses on the technical practice aspects of problem formation, model design assumption, model evaluation, and limitation analysis. This middle level critical reflection is based on our prior high level critical reflections in the AI for Just Work paper.
Low level technical practice
Low level technical practice refers to the micro scope of the specific decisions and in technical conduct. Criticality in this level of technical practice concerns about, for example, whether the use of a common evaluation metrics is scientific rigor and responsible, the scientific flaws in a conventional formed practice. Examples of critical reflections of the low level technical practice can be:
- In the paper, Metrics reloaded: recommendations for image analysis validation, Maier-Hein et al. critically analyze the pros and cons of evaluation metrics for imaging tasks.
- The paper Are We Learning Yet? A Meta Review of Evaluation Failures Across Machine Learning provides critique to the machine learning evaluation paradigm regarding internal and external validity.
- In my XAI plausibility case above, we conduct a critical examination of the plausibility metric for XAI algorithm evaluation and optimization, and identify this evaluation practice as being unscientific and unethical.
The anticipated resistance
The above different levels of approach hopefully can inspire you to think about how you can incorporate critical thinking and reflection into your current technical work to strengthen the technical immune system. In fact, I think medical image analysis is the ideal field to conduct critical technical practice, because the consequences of not conducting responsible and critical technical practice are huge. Despite the benefits of critical technical practice in the long run, currently, conducting critical technical practice is anticipated to receive many resistance and rejections, because the current technical culture priorities technical optimism (which emphasizes benefits and strengths of techniques) over scientific skepticism (which emphasizes weaknesses, limitations, flaws, and failure modes of techniques). Conducting critical technical practice may also put technical practitioners and researchers in a relatively vulnerable position. A well-known case is Timnit Gebru got fired from Google. That’s why we need collective efforts to support each other along the way. Knowing that resistance is normal in critical technical practice, we will be less likely to blame ourselves for failures and rejections. It also allows us to adjust expectations and shift our reward system from short-term external incentives, such as getting paper published, to long-term internal motivations, such as personal growth and the sense of feeling connected and supported towards a common goal.
Actionable ideas for critical technical practice
If you feel motivated to conduct critical technical practice in your work, here are a few actionable ideas for you to consider:
- Do not easily let go of your doubts and skepticism during your work. Write them down and see if you can investigate them. Imagine yourself as a detective, and these doubts can be your initial clue to discover the hidden truth in technical work.
- From your skepticism, analyze its underlying assumptions and rationales for these assumptions, question the justifications, inspect the overall framing of the question, and identify the underlying philosophical worldviews and encoded values.
- Consider organizing a reading group or working group in your technical circle, to read more about related works on critical technical practice, discuss with peers, build your local critical technical practice community and support group, and see if you can collaboratively work something out to improve the scientific rigor, responsibility, and ethics of the technical work by understanding the root problems.
Enjoy Reading This Article?
Here are some more articles you might like to read next: