<?xml version="1.0" encoding="UTF-8"?>

<?xml-stylesheet type="text/xsl" href="/static/oaitohtml.xsl"?>

<!--
<?xml-stylesheet type="text/xsl" href="/oaitohtml.xsl"?>
-->

<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
    <responseDate>2026-10-10T15:05:38Z</responseDate>
    <request verb="GetRecord" metadataPrefix="oai_dc" identifier="10.57760/sciencedb.14727" >https://www.scidb.cn/oai</request>
<GetRecord>
    <record>
    <header >
    <identifier>10.57760/sciencedb.14727</identifier>
    <datestamp>2024-04-16T14:59:58Z</datestamp>
</header>
    <metadata>
        
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:date>2024-04-16</dc:date>
  <dc:title>CPIA Dataset_Part05: A Comprehensive Pathological Image Analysis Dataset for Self-supervised Learning Pre-training</dc:title>
  <dc:identifier>doi:10.57760/sciencedb.14727</dc:identifier>
  <dc:language>en</dc:language>
  <dc:description>Pathological image analysis is a crucial field in computer-aided diagnosis. Transfer learning using models initialized on natural images has improved the downstream pathological performance. However, the lack of sophisticated domain-specific pathological initialization hinders their potential. Self-supervised learning (SSL) enables pre-training without sample-level labels, overcoming the challenge of expensive annotations. Thus, this field calls for a comprehensive dataset, similar to the ImageNet in computer vision. This paper presents a large-scale comprehensive pathological image analysis (CPIA) dataset for SSL pre-training. The CPIA dataset contains 148,962,579 images, covering over 48 organs/tissues and approximately 100 kinds of diseases, which includes two main data types: whole slide images (WSIs) and characteristic regions of interest (ROIs). And we establish a multi-scale pathological data processing workflow, combined with the diagnosis habits of senior pathologists. The CPIA dataset facilitates a comprehensive pathological understanding and enables pattern discovery explorations. Additionally, to launch the CPIA dataset, several state-of-the-art (SOTA) baselines of SSL pre-training and downstream evaluation are specially conducted.&amp;nbsp;This is the Part05 of CPIA dataset, including the CPIA-Mini and partial CPIA dataset.&amp;nbsp;The related&amp;nbsp;code&amp;nbsp;and information&amp;nbsp;are available at https://github.com/zhanglab2021/CPIA_Dataset.</dc:description>
  <dc:subject>Pathological images; Pre-training; Self-supervised learning; Large-scale dataset</dc:subject>
  <dc:creator>Nan Ying</dc:creator>
  <dc:creator>Yanli Lei</dc:creator>
  <dc:creator>Tianyi Zhang</dc:creator>
  <dc:creator>Shangqing Lyu</dc:creator>
  <dc:creator>Sicheng Chen</dc:creator>
  <dc:creator>Zeyu Liu</dc:creator>
  <dc:creator>Yu Zhao</dc:creator>
  <dc:creator>Yunlu Feng</dc:creator>
  <dc:creator>Guanglei Zhang</dc:creator>
  <dc:rights>PUBLIC</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
  <dc:type>dataset</dc:type>
  <dc:publisher>Science Data Bank</dc:publisher>
</oai_dc:dc>

    </metadata>
</record>
</GetRecord>
</OAI-PMH>