<?xml version="1.0" encoding="UTF-8"?>

<?xml-stylesheet type="text/xsl" href="/static/oaitohtml.xsl"?>

<!--
<?xml-stylesheet type="text/xsl" href="/oaitohtml.xsl"?>
-->

<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
    <responseDate>2026-10-11T03:11:50Z</responseDate>
    <request verb="GetRecord" metadataPrefix="oai_dc" identifier="10.57760/sciencedb.013ss" >https://www.scidb.cn/oai</request>
<GetRecord>
    <record>
    <header >
    <identifier>10.57760/sciencedb.013ss</identifier>
    <datestamp>2026-09-28T16:57:57Z</datestamp>
</header>
    <metadata>
        
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:date>2026-09-28</dc:date>
  <dc:title>Replication Package: Self-Reported Test-Taking Effort and Behavioral Engagement in PISA 2025 (PISA 2025 LDW)</dc:title>
  <dc:identifier>doi:10.57760/sciencedb.013ss</dc:identifier>
  <dc:language>en</dc:language>
  <dc:description>本数据集为一篇 PISA 2025 二次分析论文的复现包（replication package），包含全部分析代码与补充材料表格。论文考察 PISA 2025&amp;quot;数字世界中的学习&amp;quot;（LDW）测评中学生自我报告的努力程度与过程数据行为投入之间的关系，覆盖 84 个经济体、226,335 名具有过程数据的学生。数据产生过程：原始数据来自 OECD 于 2026 年 9 月 8 日公开发布的 PISA 2025 公共使用文件（PUF，学生问卷文件与 LDW 过程数据文件），按 OECD 服务条款注册获取。分析使用 R 4.6.1 编写的清洗与估计管线：先用 20260922_clean_puf.R 完成清洗与合并（生成 755,721 行学生级数据，其中 226,335 人具有 LDW 过程数据，占 LDW 参与者的 31.4%），再用 12_build_aggregates.R 将单元级行为记录聚合为学生级指标（每单元平均用时 TT、每单元平均动作数 A、放弃编码 code-9 与未触达单元计数等）。02_validate_volume1.R 将管线与 OECD《PISA 2025 Results (Volume I)》官方发表数字对照核验：91 个经济体的科学素养均值与官方值之差不超过 0.003 分，CMPS 均值之差不超过 0.0023 分，跨域相关之差不超过 0.00001。其余脚本在验证过的估计引擎上产出论文全部统计量：加权相关与加权回归采用最终学生权重 W_FSTUWT，标准误由 80 次 BRR-Fay（&amp;epsilon;=0.5）重复方差估计结合 10 组合理值（plausible values）按 Rubin 规则合成（总方差 = 平均抽样方差 + (1+1/10)&amp;times;插补方差）。时间与空间信息：数据对应 PISA 2025 测评周期（结果于 2026 年 9 月 8 日发布），空间覆盖 84 个参测经济体；单元级行为指标的时间分辨率为 LDW 各单元的作答时长（毫秒级记录，分析中转换为秒），空间分辨率为经济体层级。文件内容：压缩包含两个目录。code/ 目录含 10 个 R 脚本与英文 README.md（内含运行顺序、OECD 数据下载说明与预期输出），按 README 所列顺序运行即可完整复现论文数字。supplementary/ 目录含 5 个 CSV 文件，对应论文附表 S1&amp;ndash;S5：s1_descriptives_by_country.csv（252 个数据行 = 84 个经济体 &amp;times; 3 行，分别为有效 N、均值、标准差；8 列依次为经济体代码 CNT、自报努力 EFFORT1 与 EFFORT2（1&amp;ndash;10 分量表）、每单元平均用时 TT_mean_s（秒/单元）、每单元平均动作数 A_mean（次/单元）、CMPS 表现（PISA 量表分，国际均值 500、标准差 100）、家庭经济社会文化地位指数 ESCS（OECD 标准化指数）、code-9 放弃单元数 n_code9（个/学生））；s2_within_country_correlations.csv（336 个数据行 = 84 经济体 &amp;times; 4 个预测变量），记录各国 CMPS 与四个指标的加权相关系数 est 及标准误 se（无量纲）；s3_within_country_effort1_coefficients.csv（84 个数据行），为逐国回归中 EFFORT1 的系数 beta（单位：CMPS 分/1 分自评努力）及其标准误、t 值与显著性分类 class（negative_sig/negative_ns/positive_sig/positive_ns/not_estimated）；s4_within_country_by_band.csv（252 个数据行 = 84 经济体 &amp;times; 3 个能力段），记录低/中/高能力段内 CMPS&amp;times;TT、EFFORT1&amp;times;TT、EFFORT1&amp;times;CMPS 的相关系数及标准误；s5_country_codebook.csv（91 个数据行），为 ISO-3 代码与官方显示名对照，并标记 Volume I 脚注国家与是否具有过程数据。缺失情况：s3 中哥斯达黎加（CRI）与萨尔瓦多（SLV）因 ESCS 全缺失无法估计，beta/se/t 为空，class 标记为 not_estimated；s1 的 ESCS 列在部分经济体存在缺失（CRI、SLV 全缺失，美国、亚美尼亚、黎巴嫩分别缺失约 45%、42%、29%），相应单元格为空。除 OECD 自身的合理值插补外，分析未做任何插补，未剔除任何异常值。误差说明：所有统计量均按 PISA 官方方法计入复杂抽样方差（BRR-Fay）与测量不确定性（合理值插补方差），各表中的 se 列即相应标准误；未引入其他已知误差来源。This dataset is a replication package for a secondary analysis of PISA 2025. It contains the full analysis code and supplementary tables. The paper examines the relationship between students&amp;rsquo; self-reported effort and behavioral engagement from process data in the PISA 2025 &amp;ldquo;Learning in the Digital World&amp;rdquo; (LDW) assessment, covering 84 economies and 226,335 students with process data.**Data provenance.** The raw data come from the OECD PISA 2025 Public Use Files (PUF)&amp;mdash;the student questionnaire file and the LDW process-data file&amp;mdash;publicly released on 8 September 2026 and obtained under the OECD Terms of Service upon registration. Analysis uses an R 4.6.1 cleaning and estimation pipeline: `20260922_clean_puf.R` first cleans and merges the files (yielding 755,721 student-level records, of which 226,335 have LDW process data, 31.4% of LDW participants); `12_build_aggregates.R` then aggregates item-level behavioral records into student-level indicators (mean time on task per item, TT; mean number of actions per item, A; abandoned-item counts coded as code-9; and unreached-item counts). `02_validate_volume1.R` checks the pipeline against official figures in OECD *PISA 2025 Results (Volume I)*: science proficiency means for 91 economies differ from the official values by no more than 0.003 points, CMPS means by no more than 0.0023 points, and cross-domain correlations by no more than 0.00001. All remaining scripts produce the paper&amp;rsquo;s statistics on this validated estimation engine. Weighted correlations and weighted regressions use the final student weight `W_FSTUWT`; standard errors are obtained from 80 BRR-Fay (&amp;epsilon; = 0.5) replicate weights combined with 10 plausible values under Rubin&amp;rsquo;s rules (total variance = mean sampling variance + (1 + 1/10) &amp;times; imputation variance).**Time and space reference.** Data correspond to the PISA 2025 assessment cycle (results released 8 September 2026) and cover 84 participating economies. Item-level behavioral indicators are timed at the resolution of each LDW item&amp;rsquo;s response duration (recorded in milliseconds and converted to seconds for analysis), with spatial resolution at the economy level.**File contents.** The archive contains two directories. `code/` holds 10 R scripts and an English `README.md` (with run order, OECD data-download instructions, and expected outputs); running the scripts in the order listed in the README fully reproduces the paper&amp;rsquo;s numbers. `supplementary/` holds 5 CSV files corresponding to Tables S1&amp;ndash;S5: `s1_descriptives_by_country.csv` (252 data rows = 84 economies &amp;times; 3 rows: valid *N*, mean, and standard deviation; 8 columns: economy code CNT; self-reported effort EFFORT1 and EFFORT1/EFFORT2 on 1&amp;ndash;10 scales; mean time per item TT_mean_s (s/item); mean actions per item A_mean (actions/item); CMPS performance (PISA scale, international mean 500, SD 100); family socioeconomic status index ESCS (OECD-standardized index); and abandoned-item count n_code9 (items/student)); `s2_within_country_correlations.csv` (336 data rows = 84 economies &amp;times; 4 predictors), economy-level weighted correlations `est` and standard errors `se` between CMPS and each of the four indicators (dimensionless); `s3_within_country_effort1_coefficients.csv` (84 data rows), economy-specific regressions with EFFORT1 coefficient `beta` (CMPS points per 1-point self-reported effort), standard error, *t*-value, and significance class (`negative_sig` / `negative_ns` / `positive_sig` / `positive_ns` / `not_estimated`); `s4_within_country_by_band.csv` (252 data rows = 84 economies &amp;times; 3 proficiency bands), within-band correlations and standard errors for CMPS&amp;times;TT, EFFORT1&amp;times;TT, and EFFORT1&amp;times;CMPS; and `s5_country_codebook.csv` (91 data rows), ISO-3 codes with official display names, flags for Volume I footnote countries, and indicators of whether process data are available.**Missingness.** In `s3`, Costa Rica (CRI) and El Salvador (SLV) cannot be estimated because ESCS is entirely missing; `beta`, `se`, and `t` are empty and `class` is coded `not_estimated`. In `s1`, the ESCS column is missing in several economies (entirely missing for CRI and SLV; approximately 45%, 42%, and 29% missing for the United States, Armenia, and Lebanon, respectively); the corresponding cells are empty. Apart from OECD&amp;rsquo;s own plausible-value imputation, no additional imputation was performed and no outliers were removed.**Error note.** All statistics incorporate complex-survey variance (BRR-Fay) and measurement uncertainty (plausible-value imputation variance) according to official PISA methods; the `se` columns in the tables give the corresponding standard errors. No other known sources of error were introduced.</dc:description>
  <dc:subject>process data; PISA; replication; survey analysis; test-taking effort</dc:subject>
  <dc:creator>孙翊超 Yichao Sun</dc:creator>
  <dc:rights>PUBLIC</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
  <dc:type>dataset</dc:type>
  <dc:publisher>Science Data Bank</dc:publisher>
</oai_dc:dc>

    </metadata>
</record>
</GetRecord>
</OAI-PMH>