<?xml version="1.0" encoding="UTF-8"?>

<?xml-stylesheet type="text/xsl" href="/static/oaitohtml.xsl"?>

<!--
<?xml-stylesheet type="text/xsl" href="/oaitohtml.xsl"?>
-->

<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
    <responseDate>2026-10-11T05:51:14Z</responseDate>
    <request verb="GetRecord" metadataPrefix="oai_dc" identifier="10.57760/sciencedb.psych.01039" >https://www.scidb.cn/oai</request>
<GetRecord>
    <record>
    <header >
    <identifier>10.57760/sciencedb.psych.01039</identifier>
    <datestamp>2026-06-12T12:31:40Z</datestamp>
</header>
    <metadata>
        
<oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:date>2026-06-12</dc:date>
  <dc:title>A Dataset of Naturalistic Reddit Comments from Self-Declared MBTI Personality Types</dc:title>
  <dc:identifier>doi:10.57760/sciencedb.psych.01039</dc:identifier>
  <dc:language>en</dc:language>
  <dc:description>This dataset is derived from the publicly available Pushshift Reddit corpus (Baumgartner et al., 2020), which archives Reddit comments since the platform&amp;rsquo;s inception in 2005. Because this study focuses on self-declared personality labels rather than temporal trends, no filtering based on comment timestamps was applied; thus the dataset does not have a specific temporal range, though most comments were posted between the 2010s and early 2020s. Regarding spatial information, Reddit users are predominantly English speakers, primarily from English-speaking countries (e.g., United States, United Kingdom, Canada), but the platform is global and no geographic tags are included in the data.The dataset was constructed in three stages.&amp;nbsp;Stage 1 &amp;ndash; User identification:&amp;nbsp;We scanned 20 personality-focused subreddits (e.g., r/mbti, r/infp, r/intj) to identify users who publicly self-declared their MBTI type via user flair. This yielded an initial cohort of 5,044 unique users with valid four-letter MBTI labels.&amp;nbsp;Stage 2 &amp;ndash; Comment extraction and filtering:&amp;nbsp;For each identified user, we retrieved their complete comment history from the Pushshift corpus. To avoid topical bias and meta-discussion about personality, we excluded all comments posted in the same 20 personality-typing subreddits (including r/mbti, all 16 MBTI type subreddits, r/enneagram, r/socionics, and r/JungianTypology). The remaining comments originated from 1,847 other subreddits covering diverse topics (technology, lifestyle, entertainment, academia, etc.).&amp;nbsp;Stage 3 &amp;ndash; Activity filtering:&amp;nbsp;To ensure statistical reliability, we retained only users with at least 20 valid comments after exclusion. The final sample comprises 2,005 users and approximately 227,000 naturalistic comments. The median number of comments per user is 56, the mean is 113, and the distribution follows a long-tail pattern.Data processing was implemented in Python. Two lexicon-based tools were used: (1) VADER (Valence Aware Dictionary and sEntiment Reasoner; Hutto &amp;amp; Gilbert, 2014) to compute a compound sentiment polarity score for each comment (range: -1 to +1), and (2) the NRC Emotion Lexicon (Mohammad &amp;amp; Turney, 2013) to count, for each comment, the frequency of words associated with eight basic emotions (joy, trust, sadness, anger, fear, surprise, anticipation, disgust). Per-comment measures were then aggregated at the user level. The final dataset is provided as a single CSV file (mbti_reddit_emotion_dataset.csv) with 2,005 rows (one per user) and 18 columns. The columns are:user_id&amp;nbsp;(anonymized identifier)isExtrovert&amp;nbsp;(binary: 1 = E, 0 = I)isSensing&amp;nbsp;(binary: 1 = S, 0 = N)isThinking&amp;nbsp;(binary: 1 = T, 0 = F)isJudging&amp;nbsp;(binary: 1 = J, 0 = P)avgSentimentPolarity&amp;nbsp;(mean VADER compound score per user; range -0.263 to 0.768; dimensionless)stdSentimentPolarity&amp;nbsp;(standard deviation of VADER compound scores per user; proxy for emotional volatility; range 0.211 to 0.761)freqJoy,&amp;nbsp;freqTrust,&amp;nbsp;freqSadness,&amp;nbsp;freqAnger,&amp;nbsp;freqFear,&amp;nbsp;freqSurprise,&amp;nbsp;freqAnticipation,&amp;nbsp;freqDisgust&amp;nbsp;(mean frequency of emotion-associated words per comment, expressed as count per comment; non-negative real numbers)commentCount&amp;nbsp;(total number of valid comments for that user; integer)All variables are complete (no missing values) for the 2,005 users, as only users with &amp;ge;20 valid comments and a declared MBTI type were retained.Regarding measurement error: In an independent validation on 200 comments (with human ratings as ground truth), VADER achieved a Pearson correlation of r = 0.68 and a three-way (positive/neutral/negative) classification accuracy of 76.5%. Most errors (approximately 12 out of 200 cases, 6%) stem from VADER&amp;rsquo;s difficulty with sarcasm, irony, and context‑dependent sentiment. Because our analyses aggregate over tens to hundreds of comments per user (mean ~113), random per‑comment errors are substantially attenuated at the user level, rendering the aggregate metrics reliable. The NRC emotion categories show moderate to high inter-correlations (r = 0.60&amp;ndash;0.85), indicating that the dimensions are not fully independent&amp;mdash;an inherent limitation of lexicon‑based approaches.The CSV file is encoded in UTF‑8 and uses commas as delimiters. It can be read by any standard data analysis software, including Python (pandas), R, Microsoft Excel, and SPSS. To replicate the statistical analyses reported in the associated paper (OLS regressions, correlation matrices, and visualizations), users may employ the&amp;nbsp;statsmodels&amp;nbsp;library in Python or the&amp;nbsp;lm()&amp;nbsp;function in R.</dc:description>
  <dc:subject>Personality types; Myers-Briggs Type Indicator ; Emotional expression; natural language processing</dc:subject>
  <dc:creator>刘兴才</dc:creator>
  <dc:rights>PUBLIC</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by-nc/4.0/</dc:rights>
  <dc:type>dataset</dc:type>
  <dc:publisher>Science Data Bank</dc:publisher>
</oai_dc:dc>

    </metadata>
</record>
</GetRecord>
</OAI-PMH>