CATEGORY

AI research and data

Research systems, evaluation methods and data workflows for evidence-led AI decisions.

ON THIS PAGE
01Research agents
02Evaluation
03Data quality

EDITOR'S SELECTION

Latest analysis

Research & Data

OpenAI says it has reached its automated research intern milestone

OpenAI reports that supervised agents can now complete meaningful parts of its internal research workflow, while emphasizing that humans still choose directions and validate results.

· Priya Raman · 9 min read
→
Research & Data

Google DeepMind pilots double-blind evaluation for a Gemini model

A new pilot uses a confidential computing environment so Google cannot inspect an evaluator's secret prompts and the evaluator cannot access proprietary model weights.

· Elena Morris · 9 min read
→
Research & Data

Google Research introduces an autonomous planetary prediction engine

The experimental Earth AI system discovers geospatial data, engineers features, trains models and generates reports from a natural-language prediction request.

· Elena Morris · 10 min read
→
Research & Data

OpenAI study separates ChatGPT gains from critical-thinking training

A randomized experiment with more than 1,000 Bocconi students found that ChatGPT access improved polished task performance while causal-reasoning training broadened ideas.

· Priya Raman · 9 min read
→
Research & Data

How to evaluate a transcription model that rewrites speech

Gemini 3.5 Transcribe can remove filler words, format text and resolve corrections. Those features require separate tests for readability, fidelity and decision-critical accuracy.

· Priya Raman · 10 min read
→
Research & Data

How to evaluate a custom AI inference chip benchmark

OpenAI's Jalapeño results provide a timely case study in separating chip metrics, system performance and real application value.

· Priya Raman · 10 min read
→
Research & Data

How to read an AI inference energy benchmark

NVIDIA's Vera Rubin claims put throughput per megawatt in the spotlight. This framework separates useful efficiency evidence from a favorable vendor headline.

· Priya Raman · 10 min read
→
Research & Data

How to build a held-out evaluation for speech recognition

A useful ASR test set must resemble production without becoming part of the optimization loop. This guide combines temporal separation, error weighting and audio-level review.

· Priya Raman · 10 min read
→
Research & Data

How to use AI for data analysis without losing the audit trail

AI can accelerate exploration, formulas and narrative summaries, but analysts still need a reproducible path from source data to every published claim.

· Elena Morris · 10 min read
→
Research & Data

Persistent worlds could expose the failures agent benchmarks miss

A persistent environment tests memory, adaptation and social consequences across time. It can improve agent evaluation only if researchers publish tasks, baselines and failure evidence.

· Elena Morris · 9 min read
→
Research & Data

Natural language processing, from tokens to modern AI systems

NLP spans classifiers, search, embeddings and large language models. This guide explains the core ideas and practical choices without treating every text problem as a chatbot problem.

· Elena Morris · 10 min read
→
Research & Data

Machine learning algorithms: how to choose the right family for a real problem

Regression, trees, clustering and neural networks solve different kinds of problems. This practical map starts with the decision, the data and the cost of being wrong—not a list of fashionable algorithms.

· Ravi Kapoor · 13 min read
→
Research & Data

Types of machine learning models: a practical guide from regression to transformers

Model types are easier to understand when organized by what they learn, what data they consume and how their output will be used. Here is a practical taxonomy for choosing and evaluating them.

· Ravi Kapoor · 12 min read
→
Research & Data

X is letting AI draft Community Notes—but humans keep the vote

The AI Note Writer API offers a revealing model for human-AI research systems: machines can find sources and propose context, while people with diverse viewpoints decide what becomes visible.

· Priya Raman · 7 min read
→
Research & Data

Long context is not the same as long-term memory

LongMemEval shows why placing more conversation into a model's window does not guarantee reliable recall. Useful memory systems must index, retrieve, update and sometimes abstain.

· Priya Raman · 8 min read
→
Audio & Voice

Hume researchers find signs of benchmark fitting in speech recognition

Tests across 11 open speech-recognition models found cases where systems reproduced benchmark-specific text even when the audio contradicted it. Leaderboard accuracy may overstate real-world transcription quality.

· Priya Raman · 9 min read
→
Features

Why AI produces confident false answers—and how to verify them

NIST calls the problem confabulation: confidently stated erroneous content. Understanding why it happens leads to better interfaces, tests and review practices.

· Elena Morris · 9 min read
→
AI Tools

How to choose an AI search tool without losing the evidence

AI search can compress research time, but fluent synthesis can hide weak retrieval. Evaluate source coverage, citation fit, freshness and reproducibility before trusting an answer.

· Elena Morris · 9 min read
→
Model Platforms

How to evaluate an AI model API before you commit

A leaderboard cannot tell you which model belongs in your product. Measure quality, latency, cost, safety and operational fit together.

· Elena Morris · 10 min read
→
Reviews

Perplexity AI review 2026: excellent research navigation, but verification is still your job

Perplexity combines live web search, multiple model options, cited answers and persistent research projects. We evaluate where that workflow saves time—and where source quality and synthesis still need human review.

· Elena Morris · 14 min read
→
AI Models

GPT-5.6 arrives across ChatGPT, Codex and the OpenAI API

OpenAI's July flagship release brings one model family to consumer chat, coding agents and the API, followed by substantial price cuts for its Luna and Terra variants.

· Ravi Kapoor · 7 min read
→
AI Agents

ChatGPT Work turns OpenAI's chatbot into a long-running work agent

ChatGPT Work can act across apps and files, break goals into steps and stay with a project for hours, pushing ChatGPT further from answer engine to execution layer.

· Linh Nguyen · 7 min read
→