After BrowseComp was brushed to 90%, Meituan LongCat launched LoHoSearch: Frontline models collectively dropped back to less than 30%
"Search Agent Evaluation Benchmark BrowseComp was quickly overwhelmed," its performance soared from 30% to 90% and gradually became ineffective. On July 17, Meituan LongCat released a new benchmark LoHoSearch, which generates difficult problems based on a Wikipedia knowledge graph containing 7.62 million entities, aiming to push the evaluation back into a high-difficulty area and reset the standard for search agent capabilities.