ITBench-AA: A New Benchmark for Evaluating AI in Enterprise IT Tasks

IBM and Artificial Analysis unveil ITBench-AA, marking a significant step in assessing AI performance in Site Reliability Engineering tasks, with frontier models scoring below 50%.

IBM and Artificial Analysis unveil ITBench-AA, marking a significant step in assessing AI performance in Site Reliability Engineering tasks, with frontier models scoring below 50%.

The cost of AI evaluations has reached a critical threshold, reshaping the landscape of who can afford to conduct them. Recent findings reveal staggering expenses associated with evaluating AI models, highlighting the complexities and inefficiencies in current benchmarking practices.

IBM Research introduces AssetOpsBench, a benchmark system designed to evaluate AI agents in complex industrial environments, enhancing performance assessment beyond traditional metrics.

OpenAI is engaging third-party contractors to upload real work documents to assess the performance of its AI models, marking a significant step in its pursuit of advanced AI capabilities.