<?xml version="1.0" encoding="UTF-8"?>

<?xml-stylesheet type="text/xsl" href="/static/oaitohtml.xsl"?>

<!--
<?xml-stylesheet type="text/xsl" href="/oaitohtml.xsl"?>
-->

<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
    <responseDate>2026-10-10T17:11:31Z</responseDate>
    <request verb="GetRecord" metadataPrefix="oai_dc" identifier="10.57760/sciencedb.41072" >https://www.scidb.cn/oai</request>
<GetRecord>
    <record>
    <header >
    <identifier>10.57760/sciencedb.41072</identifier>
    <datestamp>2026-06-25T11:06:47Z</datestamp>
</header>
    <metadata>
        
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:date>2026-06-25</dc:date>
  <dc:title>WushuQA: A Chinese Martial Arts Question-Answering Dataset for Large Language Model Fine-Tuning</dc:title>
  <dc:identifier>doi:10.57760/sciencedb.41072</dc:identifier>
  <dc:language>en</dc:language>
  <dc:description>WushuQA is a question-answering dataset designed for instruction fine-tuning of large language models (LLMs) in the domain of Chinese martial arts (武术/Wushu).Background &amp;amp; Motivation: Chinese martial arts feature a vast and intricate knowledge system with highly specialized terminology and complex lineage structures. General-purpose LLMs consistently perform poorly in this domain &amp;mdash; confusing styles, misattributing techniques, and giving vague or incorrect explanations of core concepts. The root cause is a near-total absence of high-quality martial arts text in standard training corpora.Scale: Built upon 1,056 source texts (~2.8 million Chinese characters) drawn from six authoritative sources &amp;mdash; including the Wushu Management Center of the General Administration of Sport of China, the Chinese Wushu Association, the International Wushu Federation, Wikipedia, and Baidu Baike &amp;mdash; plus one specialized martial arts website, the dataset yields 14,380 high-quality QA pairs in JSONL format. Each record retains the original source text for full traceability.Construction Pipeline: The dataset was built through an eight-stage pipeline: rule-based cleaning &amp;rarr; structured segmentation &amp;rarr; multi-level LLM question generation &amp;rarr; semantic deduplication (embedding similarity threshold 0.92) &amp;rarr; multi-model answer generation &amp;rarr; multi-model cross-review &amp;rarr; full-scale faithfulness scanning &amp;rarr; question context completion. Starting from 15,565 initially generated questions, 14,380 passed all quality filters.Coverage: Five knowledge subdomains &amp;mdash; martial arts styles, techniques &amp;amp; theory, historical figures, competition rules, and cultural heritage &amp;mdash; across six question types: factual, explanatory, comparative, enumerative, judgmental, and comprehensive.</dc:description>
  <dc:subject>WushuQA; Chinese martial arts; large language models</dc:subject>
  <dc:creator>wu cheng xue</dc:creator>
  <dc:rights>PUBLIC</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
  <dc:type>dataset</dc:type>
  <dc:publisher>Science Data Bank</dc:publisher>
</oai_dc:dc>

    </metadata>
</record>
</GetRecord>
</OAI-PMH>