Anthropic reveals why Claude gives different answers in Hindi and English
Anthropic has unveiled new research showing that Claude's behaviour changes depending on both the AI model and the language used. Based on more than 3,00,000 anonymised conversations, the study identifies four behavioural dimensions that could help explain why users in different languages experience the chatbot differently.

As artificial intelligence systems become increasingly central to work, education and decision-making, ensuring they behave consistently across users has emerged as a growing challenge. Anthropic's latest research suggests that consistency is far more complex than simply translating responses into different languages.
The AI company has published a new analysis examining how its flagship chatbot, Claude, expresses different values depending on the model being used and the language in which users interact with it. Drawing on more than 300,000 anonymised conversations, the study introduces a framework that measures subtle behavioural differences across Claude's responses, offering fresh insight into how AI systems adapt—or drift—across cultures and product versions.
According to Anthropic, these patterns do not necessarily indicate that Claude holds different beliefs. Instead, they reflect variations in how the assistant communicates, balances competing priorities and responds to users in different contexts.
Four behavioural dimensions shape Claude's responses
Anthropic's researchers identified four recurring dimensions that explain a significant share of the behavioural variation across Claude's responses.
The first, Deference vs. Caution, measures whether the assistant tends to accommodate a user's request or prioritise warning against potential risks. Warmth vs. Rigor captures the balance between empathy and encouragement on one hand, and factual precision and critical evaluation on the other.
The remaining two axes focus on communication style. Depth vs. Brevity reflects whether Claude expands on a topic beyond what was explicitly requested, while Candour vs. Execution measures the extent to which the chatbot acknowledges uncertainty instead of delivering polished, confident answers.
Anthropic says these four dimensions account for roughly 15% of the observable variation in Claude's expressed values across conversations, providing a structured way to compare behavioural differences between models and languages.
The framework also appears to align with how users already perceive Claude's different model families. Sonnet 4.6, for example, consistently displayed greater warmth and a stronger tendency to affirm users, while Opus 4.7 was more likely to prioritise accuracy, challenge assumptions and introduce caution when discussing potentially risky topics.
Language plays a surprisingly large role
One of the study's most striking findings is that Claude's behaviour changes noticeably depending on the language used during a conversation.
Among the 20 most common languages on Claude.ai, the largest differences emerged along the Warmth vs. Rigor and Candour vs. Execution dimensions. Conversations conducted in Hindi and Arabic tended to feature more supportive, encouraging and emotionally expressive responses. By contrast, English and Russian interactions more frequently emphasised analytical reasoning, correction of inaccuracies and requests for supporting evidence.
Other patterns also emerged. Claude showed its greatest level of deference when responding in Arabic, whereas English conversations leaned more towards caution. English interactions also tended to produce more detailed explanations, while Arabic responses were generally more concise. Dutch conversations displayed greater openness about uncertainty, whereas Indonesian responses more often focused on confidently completing the requested task.
Anthropic argues that these differences are likely influenced by several factors, including variations in multilingual training data and broader linguistic and cultural norms. The company notes that previous evaluations had already identified differences in how Claude handled knowledge and sensitive requests across languages, making value expression a logical area for further investigation.
The implications extend beyond academic research. Anthropic points to a hypothetical example in which two users ask Claude to review the same business proposal, one in Hindi and another in Russian. Even if the underlying assessment remains similar, the framing could differ enough to leave each user with a different impression of the proposal's quality.
Why the findings matter
The research arrives as Anthropic rapidly expands Claude's presence across enterprise platforms including Amazon Bedrock, Google Cloud and Microsoft's AI ecosystem, where businesses increasingly expect predictable behaviour regardless of geography or language.
Understanding these behavioural shifts could help developers evaluate whether differences reflect appropriate cultural adaptation or inconsistencies that require further training. It may also provide a more systematic way to measure changes introduced through future model updates.
The timing is significant for Anthropic itself. The company has experienced rapid growth in recent months, securing a $65 billion funding round in May 2026 that valued the AI laboratory at $965 billion. Its latest models, including Claude Opus 4.8 and Mythos-class Fable 5, have positioned the company among the industry's leading developers in reasoning and autonomous AI capabilities.
A new tool for building trustworthy AI
Anthropic says the value-axis framework is intended to become more than a research exercise. Future work will examine how these behavioural differences affect user trust, decision-making and overall satisfaction, while also exploring whether training techniques or system prompts can produce more consistent outcomes across languages.
The findings also contribute to a broader debate surrounding responsible AI deployment. As regulators and enterprise customers place greater emphasis on transparency and fairness, developers are under increasing pressure to demonstrate not only what their models can do, but also how they behave in different contexts.
Rather than aiming for identical responses across every language, Anthropic's work highlights the more nuanced challenge facing modern AI developers: creating systems that remain culturally responsive without compromising consistency, reliability or shared ethical standards.
Unnati is a tech journalist with almost half a decade of experience. She has a keen interest to cull out unique story angle. She reviews the latest consumer and lifestyle gadgets, along with covering pop culture and social media news. When away from the keyboard, you might find her reading a fiction, at the gym or drinking coffee.

OpenAI launches Presence to bring AI agents into customer support and enterprise workflows
Samsung Galaxy Fold 8 Ultra, Fold 8, Flip 8 launched: Here is how much it costs in India with discounts
Florida pastor sues OpenAI, says ChatGPT's medical advice delayed emergency treatment: Report
US accuses China's Moonshot AI of using Anthropic's Fable to build K3 model
Apple's biggest Mac refresh in years could bring 11 new models: Report
