HaiAI123

Curated Global AI Tools Directory

NewHot story
94
heat index
New

Meituan releases MineExplorer, a long-horizon multimodal benchmark

10/03/2026 — 10/03, 11:12·1 sources·1 reports

Story overview

On October 3, 2026, Meituan's technical team announced that the LongCat team at Meituan has built MineExplorer, which it calls the first benchmark for minute-level long-horizon tasks in open-world settings. According to the announcement, the benchmark is meant to systematically evaluate the real capabilities of multimodal large models on tasks that require long-horizon planning and that come with hidden prerequisites — conditions a model has to uncover rather than being handed upfront.

Two aspects of the framing stand out in the announcement. First, the tasks are described as long-horizon and as running at a minute-level timescale in an open world, which places them beyond short, tightly scoped problems. Second, the evaluation target is specifically multimodal models, and the hidden prerequisites are treated as part of what the benchmark measures.

Meituan's technical team says the evaluation shows that top multimodal models exhibit a capability gap on this class of tasks, one it describes as previously overlooked. The announcement does not name the models that took part, does not give the number of tasks or how they are scored, and provides no numeric results.

That is where the matter currently stands: MineExplorer has been introduced by its builders as an evaluation benchmark, together with the finding that leading multimodal models show a previously overlooked gap on such tasks. No further details about the benchmark itself, and no follow-up steps, have been disclosed in the material available so far.

AI-generated from 1 reports · updated 2 hours ago

Latest turnMeituan's LongCat team has built MineExplorer, described as the first benchmark for minute-level long-horizon tasks in an open world, testing multimodal LLMs on long-range planning with hidden preconditions. The results point to a capability gap in top multimodal models that has largely gone unnoticed.

Reports on this story headlines open the original

Today
  1. Meituan's LongCat team has built MineExplorer, described as the first benchmark for minute-level long-horizon tasks in an open world, testing multimodal LLMs on long-range planning with hidden preconditions. The results point to a capability gap in top multimodal models that has largely gone unnoticed.

    美团技术团队AI score 75

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to the hot board →