Key takeaways
- Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning.
- Unlike standard testing methods, PerceptionBench breaks vision down into ten atomic sub-skills instead of lumping perception, knowledge…
- The 42 open-source benchmarks they analyzed show little overlap in their error profiles, so each one covers a different subset of visual…
What happened
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. 6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage. The team behind the Chinese AI assistant Kimi has introduced PerceptionBench, a benchmark that isolates and tests the visual perception of multimodal language models.
Models with nearly identical aggregate scores diverge sharply in individual categories. On average, "hallucination" is the weakest skill across the board. 6 percent. " The authors argue that many multimodal model failures typically chalked up to "reasoning errors" actually happen at the perception level. When a model botches a multi-step task, the first step, correctly reading the image, has often already gone wrong.
The same research team already released WorldVQA, a benchmark that separates object recognition from reasoning. 4 percent, and all models systematically overestimated their own confidence. A separate study by Chinese institutions, with Moonshot AI's involvement, used the BabyVision benchmark to show that frontier models fail at basic visual tasks tied to early childhood development, such as tracing lines or counting hidden blocks. 1 percent.
The researchers attribute this gap to a verbalization bottleneck where visual information gets translated into language and loses fidelity.
Why it matters
Unlike standard testing methods, PerceptionBench breaks vision down into ten atomic sub-skills instead of lumping perception, knowledge, and reasoning into a single task. Every question can be answered just by looking at the image, with no reasoning or outside knowledge required. The authors explain their approach by pointing out that existing benchmarks each capture only a narrow slice of perception errors.
The 42 open-source benchmarks they analyzed show little overlap in their error profiles, so each one covers a different subset of visual weaknesses. No single test or small group of tests was enough to capture visual perception as a whole. Rather than defining categories up front, the authors built their taxonomy from actual model errors and traced each one back to the earliest failed step in existing benchmarks.
The result is ten "skill domains": Visual Relation, Counting, Attributes, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. From an internal pool of over 17,000 verified questions, Moonshot AI is publishing 3,000 tasks. Sixty percent were derived from attributed model errors, while 40 percent were reformulated using augmented images.
The tasks seem trivial on the surface: figuring out where a symbol sits on a clock face, counting flowers inside a red box, or deciding which of two pencil cups shows a gray-pink combo versus a solid pink with a cartoon design. 6 Sol. 2 percent. 8 percent. 5 percent) trail far behind. The category-level results are more telling than the overall ranking.
What to watch
PerceptionBench breaks those questions into perception-only sub-questions, making it possible to pinpoint which specific visual ability is failing. The dataset and evaluation code are available on GitHub at MoonshotAI/PerceptionBench. 6 Sol to within a few points on general benchmarks. K3 still lags well behind in specialized areas like offensive cybersecurity and complex math. On visual perception, K3 now performs on par with its Western rivals.




