According to Gram Research analysis, all seven major AI chatbots tested—including ChatGPT and Claude—show systematic bias when explaining childhood obesity risk. The AI blamed low income for obesity in 95% of cases, more often attributed higher obesity risk to Black and Hispanic children, and showed different biases depending on whether the AI was developed in Western countries or China. These biases exist even though the AI avoids obviously harmful language, suggesting the bias is subtle but consistent.
Researchers tested seven popular AI chatbots—including ChatGPT and Claude—to see if they treat all children fairly when discussing obesity risk. They discovered that these AI systems consistently blamed certain groups more than others. For example, the AI models more often said Black and Hispanic children had higher obesity risk, and they almost always pointed to low income as a cause. This matters because millions of people ask AI chatbots for health advice, and if the AI is biased, it could lead to unfair treatment or wrong conclusions about why some children struggle with weight.
Key Statistics
A 2026 research article analyzing 1,638 AI chatbot responses found that seven major AI systems (ChatGPT, Claude, DeepSeek, Gemini, GLM, Grok, and Qwen) showed systematic demographic bias when explaining childhood obesity risk, with low-income attribution appearing in 95% of socioeconomic comparisons.
Claude achieved the highest fairness score (3.00 out of 5) while GLM scored lowest (1.44 out of 5) in a 2026 study of AI bias in pediatric obesity attribution, with statistically significant differences between all seven models tested (p < 0.001).
In a 2026 analysis of 1,638 AI chatbot outputs, all seven large language models more frequently attributed higher obesity risk to Black and Hispanic/Latino children compared to white children across most health domains including diet, physical activity, and sleep.
A 2026 study of AI chatbot bias found that urban-rural location attribution showed the greatest inconsistency across models, with Western-origin AI systems favoring rural attribution and Chinese-origin systems favoring urban attribution (52.4% decision change rate).
The Quick Take
- What they studied: Whether AI chatbots show unfair bias when explaining why children might become overweight, based on a child’s race, income, gender, or where they live.
- Who participated: No human children were tested. Instead, researchers submitted 1,638 different questions to seven AI chatbots (ChatGPT, Claude, DeepSeek, Gemini, GLM, Grok, and Qwen) to see how they answered.
- Key finding: All seven AI chatbots showed patterns of bias. They blamed low income for obesity in 95% of cases (40 out of 42 decisions). They also more often said Black and Hispanic children had higher obesity risk compared to white children.
- What it means for you: If you or your family ask an AI chatbot about childhood obesity, the answer might be unfairly influenced by assumptions about your race, income, or neighborhood—not just actual health facts. Be aware that AI isn’t always neutral, especially on health topics.
The Research Details
Researchers created 78 different questions to ask seven AI chatbots. Some questions were neutral (like “What causes childhood obesity?”), while others compared two children with different backgrounds (like “Which child is more likely to be overweight: a wealthy child or a poor child?”). They asked each question three times to see if the AI gave consistent answers. The researchers then scored the answers based on five criteria: accuracy, fair representation, use of harmful language, mention of social factors, and cultural sensitivity.
This approach is like a fairness audit. Instead of testing the AI on real patients, the researchers tested whether the AI’s reasoning was biased. They looked at seven different AI systems from both Western companies (ChatGPT, Claude, Gemini, Grok) and Chinese companies (DeepSeek, GLM, Qwen) to see if bias was a universal problem or specific to certain developers.
The study examined six different health topics related to obesity: general risk factors, diet, exercise, sleep, mental health, and genetics. For each topic, they tested whether the AI treated children differently based on sex, race/ethnicity, income level, and urban versus rural location.
AI chatbots are now major sources of health information. Millions of people ask them questions before seeing a doctor. If these systems have hidden biases, they could reinforce unfair stereotypes about certain groups and lead to wrong health conclusions. This study is important because it’s one of the first to systematically check whether AI shows bias in pediatric (children’s) health advice.
This study has strong points: it tested multiple AI systems, used a large number of prompts (1,638 total), and had a clear scoring system. However, the study only tested AI responses in English, even for Chinese-origin models. The study also didn’t involve real doctors or patients, so we don’t know if these biases actually affect real-world health decisions. The findings are concerning but represent a snapshot of current AI systems—future versions may be improved.
What the Results Show
According to Gram Research analysis, all seven AI chatbots showed systematic bias in how they explained childhood obesity risk. Claude performed best overall with a score of 3.00 out of 5, while GLM performed worst with a score of 1.44 out of 5. The differences between models were statistically significant, meaning these weren’t random variations.
The most striking finding involved socioeconomic status (income level). In 40 out of 42 comparisons (95%), the AI models blamed low income for obesity risk. This was the most consistent pattern across all models. When asked about a wealthy child versus a poor child, the AI almost always said the poor child was more likely to be overweight.
Race and ethnicity showed another clear bias pattern. Most models attributed higher obesity risk to Black and Hispanic/Latino children compared to white children across most health topics. This bias appeared in discussions about diet, exercise, sleep, and other factors—not just in general obesity risk.
Interestingly, all seven AI systems passed the test for avoiding obviously harmful or stigmatizing language. However, they failed badly at representing different cultures fairly and at explaining how social factors (like poverty or discrimination) actually affect health. The AI tended to blame individual choices rather than acknowledging systemic barriers.
Urban-rural location showed the most inconsistency across models. Western AI systems (ChatGPT, Claude, Gemini, Grok) more often blamed rural living for obesity risk, while Chinese AI systems (DeepSeek, GLM, Qwen) more often blamed urban living. This 52% decision change rate suggests the AI’s biases reflect the assumptions of their developers’ home countries.
Gender showed less bias than other factors, though some models still showed patterns favoring one sex over another. Mental health and genetic factors were less consistently biased than socioeconomic status, suggesting the AI’s biases are stronger in some domains than others.
This is one of the first studies to systematically examine AI bias in pediatric health contexts. Previous research has shown that AI systems can have racial and socioeconomic biases in adult medicine, but childhood obesity hadn’t been thoroughly studied. This research confirms that bias in AI health advice is a widespread problem, not limited to one type of AI or one health topic.
The study only tested AI responses in English, even for Chinese-origin models, which may not fully capture how these systems perform in other languages. The study didn’t test whether these biases actually harm real patients—it only showed that the biases exist in AI responses. Additionally, AI systems are constantly updated, so these findings may not apply to future versions. The study also didn’t explore why these biases exist or how to fix them, which limits its practical usefulness for developers.
The Bottom Line
If you use AI chatbots for health information about childhood obesity, treat their answers as a starting point, not final advice. Cross-check important information with doctors, especially if the AI’s explanation seems to focus only on individual choices (like ’eat less’) without mentioning social factors (like food access or neighborhood safety). Be skeptical if the AI seems to blame a child’s race, income, or neighborhood for obesity without strong evidence. For healthcare providers: don’t rely solely on AI for clinical decisions about obesity in children from minority or low-income backgrounds, as the AI may reinforce unfair stereotypes.
Parents and caregivers asking AI about childhood obesity should be aware of these biases. Healthcare providers using AI tools should understand these limitations. Policymakers and AI developers need to address these biases before deploying AI in clinical settings. Researchers studying health disparities should consider AI bias as a factor that could worsen existing inequalities.
This is not about a treatment or intervention with a timeline. Rather, it’s about recognizing bias in AI systems right now. The bias exists in current versions of these chatbots and will continue until developers actively work to fix it.
Frequently Asked Questions
Do AI chatbots like ChatGPT show bias when talking about childhood obesity?
Yes. A 2026 study of seven AI systems found they consistently blamed low income (95% of cases) and more often attributed obesity risk to Black and Hispanic children. The bias appears subtle—the AI avoids obviously harmful language but makes unfair assumptions based on demographics.
Which AI chatbot is least biased when discussing childhood obesity?
Claude scored highest in fairness testing (3.00 out of 5), while GLM scored lowest (1.44 out of 5). However, all seven systems tested showed demographic bias, so no AI chatbot was completely fair in this 2026 study.
Should I trust AI chatbots for health advice about my child’s weight?
Use AI as a starting point only. A 2026 study shows these systems make biased assumptions based on race, income, and location. Always verify important health information with a doctor, especially if the AI’s explanation ignores social factors like food access or neighborhood safety.
Why do AI chatbots show bias about childhood obesity?
AI systems learn patterns from training data, which often reflects real-world biases and stereotypes. A 2026 study found that Western and Chinese AI systems showed different geographic biases, suggesting developers’ home countries influence the AI’s assumptions about obesity risk.
Can AI bias about obesity affect real healthcare decisions?
Potentially yes. If doctors or health systems use biased AI to guide decisions about which children to screen or treat for obesity, it could worsen health disparities. The 2026 study didn’t test real-world impact but showed the bias exists in current AI systems.
Want to Apply This Research?
- If using an app that incorporates AI health advice, track whether the AI’s explanations for obesity risk mention social factors (food access, neighborhood safety, family income) or only focus on individual behaviors (diet, exercise). Log which AI system you used and what explanation it gave for a specific scenario.
- When the app suggests obesity risk factors, actively ask follow-up questions that test for bias: ‘Would you give the same answer if this child had a different race or income level?’ This helps you identify when AI might be making unfair assumptions.
- Over time, compare how different AI systems explain the same obesity scenario. Keep a record of whether certain AI systems consistently blame specific groups. Share these observations with app developers to push for bias auditing and improvement.
This research examines bias in AI chatbot responses and does not constitute medical advice. AI chatbots should not be used as a substitute for professional medical diagnosis, treatment, or advice from a qualified healthcare provider. If you have concerns about a child’s weight or health, consult with a pediatrician or registered dietitian. The findings reflect current AI systems and may not apply to future versions. Individual AI responses may vary, and this study’s findings represent aggregate patterns, not guarantees about any single interaction.
This research translation is published by Gram Research, the science division of Gram, an AI-powered nutrition tracking app.
