Peilin Zhou, Bruce Leon, Xiang Ying, Can Zhang, Yifan Shao, Qichen Ye, Dading Chong, Z. Jin et al.
A high-difficulty benchmark designed to evaluate Chinese web browsing ability, comprehensively measuring LLM agents' search and reasoning capabilities through 289 multi-hop questions.
Existing benchmarks like BrowseComp focus on English and fail to reflect the linguistic, infrastructural, and censorship-related complexities of the Chinese web. There is a lack of benchmarks to evaluate LLMs' actual browsing ability in the Chinese web environment.
Each question is reverse-engineered to have a short, objective, and verifiable answer (e.g., date, number, proper noun). A two-stage quality control protocol ensures high difficulty and answer uniqueness. Over 20 state-of-the-art LLMs and agentic search systems are evaluated.
Most models achieve accuracy below 10%, and the best-performing system (OpenAI DeepResearch) reaches only 42.9%. This demonstrates that effective retrieval strategies, sophisticated reasoning, and information reconciliation are required for Chinese web browsing. The dataset, construction guidelines, and benchmark results are publicly released.