Listen to this Post

Since October 2024, compar:IA has offered thousands of internet users the opportunity to evaluate and compare AI model outputs blindly, choosing their preferred response without knowing the model behind it. This innovative approach has generated a publicly available dataset, forming the foundation of the first Francophone participatory ranking of AI models, developed in collaboration with the PEReN (Pôle d’expertise de la régulation numérique). The goal of the ranking is not to declare a single “best” model, but to provide transparency and insight into the evolving ecosystem of generative AI based on real user preferences. The full ranking can be accessed at comparia.beta.gouv.fr/ranking and on Hugging Face.
A Transparent and Participatory Ranking
The compar:IA ranking relies on votes collected since the platform’s public launch in October 2024 and is updated weekly. Each vote is a direct duel between two models, with the preferred response winning. What sets this ranking apart is its openness and reproducibility: all vote data is published under the Etalab 2.0 license, and the methodology for calculating scores is fully transparent. Researchers and enthusiasts can access the data and notebooks on GitHub and Hugging Face to verify or reuse the results.
Importantly, this ranking does not measure technical performance or factual accuracy. It captures user preferences at a given moment and can change over time as new models emerge or user demographics evolve. Its main aims are to illuminate trends in the generative AI ecosystem, encourage model diversity—including open-source options—and integrate new evaluation criteria such as energy efficiency.
Trends Observed (October 2025)
The October 23, 2025 ranking included 60 models from 15 different developers, reflecting a multipolar, globally diverse ecosystem. Major contributors include Mistral AI (10 models), OpenAI (10), Google (9), Meta (6), and Alibaba (5). Emerging players like xAI, 01.ai, and Liquid illustrate rapid market expansion. Notably, open-weight models account for 67% of the ranking, highlighting the growing quality and influence of the open-source AI community.
Energy consumption also emerges as a key factor. Using Ecologits’ GenAI Impact methodology, models’ energy use per 1,000 tokens ranges from 3–7 Wh for the most efficient to 238 Wh for the least. Compact models like Gemma 3-12B and Qwen 3-32B demonstrate that strong perceived performance can coincide with energy efficiency, challenging assumptions that larger models are inherently better.
Transparent and Reproducible Methodology
Compar:IA employs the Bradley-Terry model to convert binary duel votes into a probabilistic ranking, with 95% confidence intervals calculated via bootstrap methods. Users’ votes create a “duel matrix” that feeds into score calculations, allowing anyone to reproduce results using publicly available data. This methodology ensures that the ranking reflects both user preferences and the uncertainty associated with each model. The platform initially tested the Elo system but opted for Bradley-Terry as AI models’ performance remains static once deployed.
Limitations and Future Directions
The ranking reflects the preferences of compar:IA users, without socio-demographic profiling, making it a non-representative but collective snapshot. Preferences are subjective, influenced by style, brevity, or clarity, and binary votes limit nuance. Compar:IA complements technical and factual evaluations rather than replacing them.
Future improvements could include thematic sub-rankings by task, complexity-based preference analysis, anonymized user profiling, expansion to European languages, and expert fact-checking to correlate user preferences with factual accuracy.
What Undercode Say:
Compar:IA represents a remarkable experiment in participatory AI evaluation, highlighting a shift toward collective intelligence and transparency in a field often dominated by opaque corporate benchmarks. The methodology demonstrates that it is possible to generate a rigorous ranking from purely subjective preferences while maintaining reproducibility and credibility.
This ranking provides insight into emerging trends: open-source models are increasingly competitive not just technically, but in user satisfaction, reflecting a community-driven ethos that prioritizes accessibility and energy efficiency. The emphasis on environmental metrics is particularly striking, as it forces the industry and users alike to consider the broader impact of AI beyond immediate performance.
Interestingly, user preferences often diverge from technical performance. Smaller, energy-efficient models are frequently favored, suggesting that user experience—clarity, conciseness, and tone—matters as much as raw capability. This signals an important lesson for AI developers: optimization should balance speed, accuracy, and user perception rather than pursuing size alone.
The ranking’s open nature fosters experimentation and engagement across the AI community. By providing raw data and reproducible scoring methods, compar:IA encourages independent verification, a level of transparency rarely seen in AI benchmarking. This approach could inspire similar frameworks globally, establishing participatory ranking as a complementary tool to conventional evaluation metrics.
Yet, challenges remain. Without user profiling, we cannot fully understand biases or the influence of task complexity on preferences. Binary voting limits subtlety, potentially masking nuanced insights. Expanding demographic data and task-specific rankings could enhance interpretability and credibility. Additionally, combining participatory rankings with expert-led factual and technical evaluations would offer a holistic view of AI model quality.
Ultimately, compar:IA represents more than a ranking: it is a lens into collective human judgment applied to artificial intelligence. By valuing user perception alongside technical metrics, it democratizes evaluation, encourages model diversity, and raises awareness of environmental impacts. It positions users as active participants in shaping AI development, rather than passive consumers of technology.
Fact Checker Results:
✅ Voting data is publicly available under Etalab 2.0 license.
✅ Energy consumption estimates are derived from open-source models using GenAI Impact methodology.
❌ Ranking does not measure factual accuracy or technical performance.
Prediction:
As participatory AI rankings like compar:IA gain traction, we are likely to see a growing influence of user perception on AI development priorities. Models that are smaller, energy-efficient, and provide pleasant, concise responses may gain prominence. 🌱💡 Integration of multi-language support and thematic sub-rankings will further diversify the AI ecosystem and empower users across regions to shape AI evolution according to their preferences.
🕵️📝✔️Let’s dive deep and fact‑check.
References:
Reported By: huggingface.co
Extra Source Hub (Possible Sources for article):
https://www.reddit.com/r/AskReddit
Wikipedia
OpenAi & Undercode AI
Image Source:
Unsplash
Undercode AI DI v2
Bing
🔐JOIN OUR CYBER WORLD [ CVE News • HackMonitor • UndercodeNews ]
📢 Follow UndercodeNews & Stay Tuned:
𝕏 formerly Twitter 🐦 | @ Threads | 🔗 Linkedin | 🦋BlueSky | 🐘Mastodon




