Creeta — AI developer tools & ecosystem news
The doctor-vs-AI health exam that only one lab graded
GPT-5.5 Instant outscored physicians on HealthBench Professional. OpenAI built the benchmark, supplied the physicians, and ran the evaluation.
Gemini Omni is paywalled. 3.5 Flash is the backend.
3.5 Flash: GA, $1.50/M input, API-callable. Gemini Omni: subscription-only, no endpoint. Decision guide for builders.
One GitHub PR is how you ship to Grok Build's marketplace
What each Grok Build plugin bundles, who the day-one partners are, and how to submit your own extension via GitHub PR.
Copilot GA'd as a standalone workspace — not as a GitHub App
GitHub Copilot's standalone app GA in 2026: worktree isolation, org admin gating, and open spending questions explained.
Firefly routes to Kling, Veo, Runway. Your IP, not Adobe's.
Firefly AI assistant routes to Kling, Veo, Runway, and 25+ other models. Adobe's indemnity covers native outputs only.
Creative Agent spans the CC suite now. Here's the paid gate.
Firefly AI Assistant (public beta, April 2026) is cross-app. This covers eligibility, invocation steps, and supervision.
Blackmail dropped from 96% to 0%. Here's the asterisk.
May 2026 alignment paper: how Anthropic cut Claude's blackmail rate from 96% to 0% and what the limits are.
Grok joins Databricks at DAIS — bring your own xAI credential
Grok 4.3 in Databricks Agent Bricks via BYOK. Unity AI Gateway controls, $5/1k tool calls, open partnership terms.
CrankGPT: Pi 5 Offline Voice AI — 0.8s TTFB, No Grid, Full Benchmark Breakdown
Pi 5, hand crank, no internet: CrankGPT's full ASR/TTS/LLM stack and llama-bench latency figures explained.
Activations into English: 4× better at surfacing hidden goals
Anthropic's NLAs map activations to English, exposing hidden goals 4× more than SAEs — and where they confabulate.
quicktok: 11× on tiktoken, author-reported, no README
quicktok: C++20 SIMD tiktoken replacement, byte-identical, 11× reported. No README. No independent reproduction.
Chronic management AI vs PCPs: 94% precision, simulated only
AMIE matched PCPs on 15 chronic management axes (94% vs 67% precision) in a Nature 2026 simulation. RxQA, Dialogue+Mx split, and key caveats.