Creeta — AI developer tools & ecosystem news

The doctor-vs-AI health exam that only one lab graded

GPT-5.5 Instant outscored physicians on HealthBench Professional. OpenAI built the benchmark, supplied the physicians, and ran the evaluation.

Gemini Omni is paywalled. 3.5 Flash is the backend.

3.5 Flash: GA, $1.50/M input, API-callable. Gemini Omni: subscription-only, no endpoint. Decision guide for builders.

One GitHub PR is how you ship to Grok Build's marketplace

What each Grok Build plugin bundles, who the day-one partners are, and how to submit your own extension via GitHub PR.

Copilot GA'd as a standalone workspace — not as a GitHub App

GitHub Copilot's standalone app GA in 2026: worktree isolation, org admin gating, and open spending questions explained.

Firefly routes to Kling, Veo, Runway. Your IP, not Adobe's.

Firefly AI assistant routes to Kling, Veo, Runway, and 25+ other models. Adobe's indemnity covers native outputs only.

Creative Agent spans the CC suite now. Here's the paid gate.

Firefly AI Assistant (public beta, April 2026) is cross-app. This covers eligibility, invocation steps, and supervision.

Blackmail dropped from 96% to 0%. Here's the asterisk.

May 2026 alignment paper: how Anthropic cut Claude's blackmail rate from 96% to 0% and what the limits are.

Grok joins Databricks at DAIS — bring your own xAI credential

Grok 4.3 in Databricks Agent Bricks via BYOK. Unity AI Gateway controls, $5/1k tool calls, open partnership terms.

CrankGPT: Pi 5 Offline Voice AI — 0.8s TTFB, No Grid, Full Benchmark Breakdown

Pi 5, hand crank, no internet: CrankGPT's full ASR/TTS/LLM stack and llama-bench latency figures explained.

Activations into English: 4× better at surfacing hidden goals

Anthropic's NLAs map activations to English, exposing hidden goals 4× more than SAEs — and where they confabulate.

quicktok: 11× on tiktoken, author-reported, no README

quicktok: C++20 SIMD tiktoken replacement, byte-identical, 11× reported. No README. No independent reproduction.

Chronic management AI vs PCPs: 94% precision, simulated only

AMIE matched PCPs on 15 chronic management axes (94% vs 67% precision) in a Nature 2026 simulation. RxQA, Dialogue+Mx split, and key caveats.

Showing of 718 posts

AI developer tools and ecosystem news for developers and technical founders

Sign up for insights and ideas

Subscribe for the latest news, stories, tips, and updates.

Subscribe