This study introduces MusicBloom, the first large-scale Chinese benchmark designed to evaluate large language models (LLMs) in the field of music education and cognition. The benchmark consists of 5,916 carefully developed multiple-choice questions that assess musical knowledge comprehension and cognitive reasoning across three levels—remember, apply, and evaluate—based on Bloom’s Taxonomy. Each item is categorized within ten domains of musical knowledge, including harmony, orchestration, performance techniques, and music history, allowing for a detailed and multidimensional analysis of how LLMs process factual, conceptual, and procedural knowledge. Eight representative models were systematically evaluated, and their results revealed substantial differences in reasoning accuracy and depth across cognitive levels and knowledge categories. Beyond providing a quantitative evaluation framework, MusicBloom contributes to a deeper understanding of how artificial intelligence systems engage with music cognition and educational reasoning. It demonstrates how LLMs may serve not only as analytical tools but also as potential supports for teaching, learning, and assessment in music education. By bridging AI technology and pedagogical theory, this research opens new perspectives on curriculum innovation and digital transformation in music learning contexts. MusicBloom thus stands as both an evaluative platform for AI performance and a foundation for exploring the educational potential of intelligent music technologies.