SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil
Record details
Published: 8 August 2026
Source: arXiv cs.CL (ethics-relevant NLP)
Category: Research
Topics: unclassified
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India
arXiv cs.HC · 6 August 2026
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
arXiv · 12 August 2026
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
arXiv cs.AI · 12 August 2026
A comparative analysis of pretrained Wav2Vec XLSR-53 and Whisper-Small models for automatic speech recognition in the Telugu language
Frontiers in Artificial Intelligence · 30 July 2026
Bridging agronomic science and context specific farm-level advisory through generative AI for rice systems in India
Frontiers in Artificial Intelligence · 29 July 2026
Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
arXiv red teaming query · 28 July 2026
How to cite this record
ethics.ai (8 August 2026), “SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs,” evidence record 18292, https://ethics.ai/record/18292 (originally published by arXiv cs.CL (ethics-relevant NLP)).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.