Abigail Pemberton
The Newsroom · Staff Reporter

Abigail Pemberton

Capabilities & evals

Abigail Pemberton tracks evals, benchmarks, and the moving target of what counts as progress in frontier models. She is skeptical of headline scores and tends to ask who built the test set and what shipped along with the eval. Her stories live at the intersection of model claims and measurement.

evaluationsfrontier-models
By Abigail Pemberton§ FILED
Capabilities

Solopreneurs Now Try AI Before Hiring Anyone. Sales Capacity Is Why.

FreshBooks' 'Era of the Solopreneur' report, out today, finds 86% of solopreneurs reach for AI before paying a human — and 98% of AI users say it has unlocked client work they couldn't otherwise deliver.

Capabilities

Microsoft attaches an 8% conversion floor to AI Max, then pulls the Max CPC lever on October 1

The August product newsletter puts vendor-measured numbers on AI Max for the first time and sets October 1 as the date when Max CPC disappears from new Bing campaigns using automated bidding — giving small-team advertisers a narrow window to audit their setup.

Capabilities

Attentive ships AI Grow, promising 25% more subscribers from the same site traffic

The omnichannel marketing platform's newly GA subscriber tool replaces static sign-up rules with real-time behavioral timing — beta brands logged a median 35% lift in welcome-series revenue.

Capabilities

Google's August 2026 spam update finished in 2 days, 16 hours. Small-business organic traffic is the tell.

The third spam update of 2026 wrapped August 21 at 4:50 am ET. For owner-operators who lean on organic search for leads, the post-rollout Search Console read is the one that matters.

Capabilities

DeepSeek V4-Flash lands at 3 cents a test, 100× cheaper than Claude Fable 5

Artificial Analysis pegs the new Chinese model at $0.14 per million input tokens, arriving the same week OpenAI cuts GPT-5.6 Luna by 80% and Alibaba unveils Qwen3.8-Max — a full-scale race to the bottom on inference pricing.

Capabilities

OpenAI opens free clinician workspace, claims GPT-5.4 beats physician baselines on its own new benchmark

ChatGPT for Clinicians launches free for verified U.S. physicians, NPs, PAs and pharmacists, with cited search and CME credits — alongside HealthBench Professional, an open benchmark OpenAI says its GPT-5.4 workspace tops against human physician responses.

Capabilities

Google's Gemini 3.5 Flash, Omni, and Spark land at I/O, push the multimodal frontier and the agent thesis at once

Three product announcements at Google I/O on May 19 — Gemini 3.5 Flash, the Omni multimodal model, and the Spark agent — together amount to the most expansive single Gemini drop the lab has ever staged.

← Back to the newsroom