<?xml version="1.0" encoding="UTF-8"?>

<?xml-stylesheet type="text/xsl" href="/static/oaitohtml.xsl"?>

<!--
<?xml-stylesheet type="text/xsl" href="/oaitohtml.xsl"?>
-->

<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
    <responseDate>2026-10-11T21:29:28Z</responseDate>
    <request verb="GetRecord" metadataPrefix="oai_dc" identifier="10.57760/sciencedb.j00001.01681" >https://www.scidb.cn/oai</request>
<GetRecord>
    <record>
    <header >
    <identifier>10.57760/sciencedb.j00001.01681</identifier>
    <datestamp>2026-07-22T15:56:48Z</datestamp>
</header>
    <metadata>
        
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:date>2026-07-22</dc:date>
  <dc:title>A Dataset for Vegetable Disease Knowledge Graph Construction</dc:title>
  <dc:identifier>doi:10.57760/sciencedb.j00001.01681</dc:identifier>
  <dc:language>en</dc:language>
  <dc:description>This dataset is aimed at the field of vegetable diseases and is mainly used for constructing a knowledge graph of vegetable diseases. The data content revolves around 31 vegetable crops and their related diseases, systematically organizing knowledge such as disease names, disease categories, symptom manifestations, affected areas, pathogens, pathogen categories, climate and environmental triggering factors, cultivation and management triggering factors, agricultural control measures, disease occurrence patterns, and pesticide control. The data is organized in the form of knowledge triplets, which can be directly mapped to entity nodes and relationship edges in the Neo4j graph database. It contains a total of 4017 entities and 9040 knowledge triplets, involving 12 types of entities and 11 types of semantic relationships. It can provide basic data support for the construction of vegetable disease knowledge graphs, knowledge retrieval, intelligent question answering, and auxiliary diagnosis applications. The dataset consists of two files, of which the main data file is vegetable_deseate_data.exe, which is composed of 1805 knowledge triplets; Neo4j_iimport.cypher is an executable Cypher script file that can be directly imported into the Neo4j graph database to build a visual vegetable disease knowledge graph.&amp;nbsp;The data in the dataset mainly comes from professional books related to vegetable diseases, authoritative domestic websites, and agricultural information platforms. In the process of data construction, firstly, OCR technology is used to recognize the content of books, and relevant unstructured text data is obtained by combining website retrieval and information collection; Subsequently, the recognition results are manually proofread to correct recognition errors and typos, and data cleaning is completed through methods such as sorting, merging, and deduplication; On this basis, relevant knowledge is extracted, summarized, and structured based on a pre designed knowledge graph pattern, and finally the data is verified through consistency checks. Through the quality control of multiple stages mentioned above, the accuracy and credibility of the dataset have been ensured.&amp;nbsp;In terms of data quality, this dataset has been cleaned and corrected as much as possible for duplicate records, inconsistent expressions, and obvious errors. However, due to the large number of original data sources and differences in expression methods, some fields may still have a small amount of missing or semantic induction bias. The missing situation is mainly reflected in the lack of certain detailed attributes in some disease records, such as specific incidence environments or prevention and control details. The errors may mainly come from differences in the content of the original data, induction bias during manual organization, and information compression caused by standardized terminology processing. Overall, this dataset can systematically reflect the basic knowledge in the field of vegetable diseases, providing a data foundation for subsequent knowledge graph construction and related research.&amp;nbsp;</dc:description>
  <dc:subject>vegetable diseases; disease control; knowledge graph; Neo4j</dc:subject>
  <dc:creator>yuan xin yi</dc:creator>
  <dc:creator>Zhang Shengqi</dc:creator>
  <dc:creator>Zhang Wu</dc:creator>
  <dc:creator>Jiang Dan</dc:creator>
  <dc:creator>Qian Yue</dc:creator>
  <dc:rights>PUBLIC</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
  <dc:type>dataset</dc:type>
  <dc:publisher>Science Data Bank</dc:publisher>
</oai_dc:dc>

    </metadata>
</record>
</GetRecord>
</OAI-PMH>