Meituan releases MineExplorer, a long-horizon multimodal benchmark
10/03/2026 — 10/03, 11:12·1 sources·1 reports
Story overview
On October 3, 2026, Meituan's technical team announced that the LongCat team at Meituan has built MineExplorer, which it calls the first benchmark for minute-level long-horizon tasks in open-world settings. According to the announcement, the benchmark is meant to systematically evaluate the real capabilities of multimodal large models on tasks that require long-horizon planning and that come with hidden prerequisites — conditions a model has to uncover rather than being handed upfront.
Two aspects of the framing stand out in the announcement. First, the tasks are described as long-horizon and as running at a minute-level timescale in an open world, which places them beyond short, tightly scoped problems. Second, the evaluation target is specifically multimodal models, and the hidden prerequisites are treated as part of what the benchmark measures.
Meituan's technical team says the evaluation shows that top multimodal models exhibit a capability gap on this class of tasks, one it describes as previously overlooked. The announcement does not name the models that took part, does not give the number of tasks or how they are scored, and provides no numeric results.
That is where the matter currently stands: MineExplorer has been introduced by its builders as an evaluation benchmark, together with the finding that leading multimodal models show a previously overlooked gap on such tasks. No further details about the benchmark itself, and no follow-up steps, have been disclosed in the material available so far.
AI-generated from 1 reports · updated 2 hours ago
Latest turnMeituan's LongCat team has built MineExplorer, described as the first benchmark for minute-level long-horizon tasks in an open world, testing multimodal LLMs on long-range planning with hidden preconditions. The results point to a capability gap in top multimodal models that has largely gone unnoticed.

Reports on this story headlines open the original
Meituan's LongCat team has built MineExplorer, described as the first benchmark for minute-level long-horizon tasks in an open world, testing multimodal LLMs on long-range planning with hidden preconditions. The results point to a capability gap in top multimodal models that has largely gone unnoticed.
美团技术团队AI score 75
Other stories people are talking about
- 554NVIDIA launches 64GB DGX Spark desktop AI computer at $4,9998 sources
- 539RisingApple tightens macOS Full Disk Access over AI agent risks7 sources
- 400Meta open-sources Muse Gadgets firmware and SDK5 sources
- 234SurgeAmazon Weighs Moving $8 Billion of Nvidia Chips into a Financing Vehicle3 sources
- 227Anthropic launches Claude Frontier Academy with $100M3 sources
- 226SurgeHugging Face open-sources AstaBrief for fast report generation3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
