<?xml version="1.0" encoding="UTF-8"?>

<?xml-stylesheet type="text/xsl" href="/static/oaitohtml.xsl"?>

<!--
<?xml-stylesheet type="text/xsl" href="/oaitohtml.xsl"?>
-->

<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
    <responseDate>2026-10-11T01:34:47Z</responseDate>
    <request verb="GetRecord" metadataPrefix="oai_dc" identifier="10.57760/sciencedb.42815" >https://www.scidb.cn/oai</request>
<GetRecord>
    <record>
    <header >
    <identifier>10.57760/sciencedb.42815</identifier>
    <datestamp>2026-07-20T17:59:42Z</datestamp>
</header>
    <metadata>
        
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:date>2026-07-20</dc:date>
  <dc:title>Constraint-Free Post-Training Quantization for Low-Bit Transformers-raw data</dc:title>
  <dc:identifier>doi:10.57760/sciencedb.42815</dc:identifier>
  <dc:language>en</dc:language>
  <dc:description>The prohibitive parameter counts and computational overhead of Transformer models hamper their deployment on resource-constrained edge devices, thereby restricting their real-world applications. Quantized low-bit Transformers demonstrate a competitive edge by requiring less storage and delivering significantly accelerated inference. However, mainstream methods are limited by strict constraints such as fixed quantization intervals, thus leading to large quantization errors (i.e., clipping and rounding errors) and substantial performance degradation. To address this issue, we propose Constraint-Free Post-training Quantization (CFQuant), which enables flexible quantization by modeling activation distributions through three key steps. First, CFQuant adaptively estimates the density of activation distributions and minimizes quantization errors through iterative search during calibration. Second, an Efficient Scale-shift Algorithm (ESA) is designed to dynamically adjust activation distributions, reducing the distribution shift between the calibration and inference stages. Moreover, Matrix Multiplication with Lookup Table (MM-LUT) is designed to accelerate inference, where floating-point multiplications are converted into pre-computed lookup operations during inference, thereby bringing substantial efficiency gains. Extensive experiments on vision, language, and multimodal tasks across various Transformermodels demonstrate the flexibility and effectiveness of CFQuant.</dc:description>
  <dc:subject>Deep learning; post-training quantization; Transformer; quantization error; self-attention</dc:subject>
  <dc:creator>jiang jia hao</dc:creator>
  <dc:creator>Yin Peng</dc:creator>
  <dc:creator>Wang Xuanhan</dc:creator>
  <dc:creator>Zeng Pengpeng</dc:creator>
  <dc:creator>Song Jingkuan</dc:creator>
  <dc:rights>PUBLIC</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
  <dc:type>dataset</dc:type>
  <dc:publisher>Science Data Bank</dc:publisher>
</oai_dc:dc>

    </metadata>
</record>
</GetRecord>
</OAI-PMH>