Zhiyong AIZhiyong AI
Editorial · public evidence · verifiable

AI changes.
What deserves attention?

Not another news recap. A readable view of what changes in real workflows, costs, and limits.

12 published postsRSSAgent JSON Feed
Latest insight · Models

After Fitting 27B Into 5.93 GB, the Bottleneck Moves

Ternary Bonsai 2 shows that ternary quantization can lower the loading barrier for a 27B model, while shifting the real competition to runtime efficiency and long-horizon reliability.

2026-09-185 min readEditorial

All insights

Sorted by latest update

Models · Analysis

After Fitting 27B Into 5.93 GB, the Bottleneck Moves

Ternary Bonsai 2 shows that ternary quantization can lower the loading barrier for a 27B model, while shifting the real competition to runtime efficiency and long-horizon reliability.

TopicTernary Bonsai 2
Developer tools · Analysis

Local Agents Are Hitting a Harness Bottleneck, Not a Model Bottleneck

The useful lesson from this open-source harness ranking is not which project has the most stars, but whether a local agent can turn its model, context, tools, and permissions into a verifiable runtime contract.

TopicAgent Harness, Ollama
Research · Analysis

The Hard Part Is Not Success, but Reproducible Success

IBM Research shows that an agent’s average score can conceal a separate reliability problem that must be measured and repaired on its own.

TopicAI agents, Agent reliability, ALTK-Evolve, Consistency Analyzer, AppWorld
Research · Analysis

Turning Papers into Tools Still Falls Short of Scientific Reproduction

Paper2Agent turns installation, execution, and validation into callable workflows, but reproducible procedures are not the same as validated scientific conclusions.

TopicPaper2Agent, Scientific reproducibility, MCP, Research agents, AlphaGenome
Agents · Analysis

Why Cheaper Tasks Can Still Produce a Larger Agent Bill

The Databricks and Steve Yegge cases show that coding-agent competition has moved beyond model pricing to the control of task boundaries, usage scale, and delivered outcomes.

TopicAI agents, Coding agents, GPT-6 Astra, Databricks, Software engineering
Infrastructure · Analysis

Video Generation Needs More Than 4-Bit Attention to Get Faster

VC-Attention tackles value quantization error and FP32 softmax in one kernel, while showing why low-bit gains depend heavily on the GPU and the full inference pipeline.

TopicVC-Attention, Video generation, Diffusion Transformer, Low-bit quantization, Softmax, GPU inference
Safety & governance · Analysis

Misalignment Is Becoming an Operations Problem

With three review tracks and six training incident reports, OpenAI is turning model anomalies from isolated research findings into a time-bound risk process.

TopicModel misalignment, Reinforcement learning, AI safety, Risk governance