Apodex today introduced TRACES, a novel benchmark designed to evaluate AI on one of the hardest challenges in artificial ...
MLPerf Client measures how effectively PCs—from laptops and desktops to workstations—run AI workloads locally. With the release of v2.0, the benchmark now includes new Agentic AI and Image Generation ...
Researchers are racing to develop more challenging, interpretable, and fair assessments of AI models that reflect real-world use cases. The stakes are high. Benchmarks are often reduced to leaderboard ...
Xiaomi's MiMo AI team has open-sourced MiMo Code V0.1.0, a terminal-native AI coding assistant that the Chinese electronics giant says outperforms Anthropic's Claude Code on key agentic coding ...
Z.ai has released GLM-5.3 with a CyberGym benchmark score of 84.5 percent, placing it ahead of rival AI models used for ...
In a new benchmark named Vibe Code Bench, OpenAI’s GPT-5.1 achieved the highest level of accuracy in completing a series of software engineering tasks, narrowly beating rival Anthropic’s Claude 4.5 ...
Secure Code Warrior, a leader in AI software governance and developer security upskilling, today introduced the SCW AI Trust Index, a living benchmark for AI coding security that grows with every new ...
China's AI developer claims GLM-5.3 rivals leading Western models in vulnerability discovery and has identified thousands of ...
Are AI benchmarks really the gold standard we’ve been led to believe? Matt Wolfe walks through how these widely accepted metrics, designed to measure the performance of artificial intelligence systems ...
AI benchmarks, often seen as the gold standard for evaluating model performance, may not be as reliable as they appear. Better Stack explores how practices like reward hacking and benchmark ...
Tech Times on MSN
Benchmark contamination detection inside AI models: New method survives RL post-training
Benchmark contamination detection method Excess Separability uses AI model activation geometry to detect whether benchmark ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results