The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014. The node label in this case is the community, or “subreddit”, that a post belongs to. 50 large communities have been sampled to build a post-to-post graph, connecting posts if the same user comments on both. In total this dataset contains 232,965 posts with an average degree of 492. The first 20 days are used for training and the remaining days for testing (with 30% used for validation). For features, off-the-shelf 300-dimensional GloVe CommonCrawl word vectors are used.
Source: https://arxiv.org/pdf/1706.02216.pdf
Image Source: https://minimaxir.com/2016/05/reddit-graph/
Variants: The Reddit Ethereum Dataset, The Reddit COVID Dataset, The Reddit Climate Change Dataset, Ten Million Reddit Answers, squadshifts reddit, Six Months of GME on Reddit, Reddit TIFU, Reddit /r/WallStreetBets data for August of 2021, Reddit /r/NoNewNormal dataset, Reddit Norm Violations, Reddit (multi-ref), REDDIT-MULTI-5k, REDDIT-MULTI-12K, Reddit Ideology Database, Reddit Engagement Dataset, Reddit C-SSRS, Reddit cryptocurrency data for August 2021, Reddit Corpus, Reddit Conversation Corpus, REDDIT-BINARY, REDDIT-B, REDDIT-5K, REDDIT-12K, Pushshift Reddit, PolyAI Reddit, One Year of Doge on Reddit, One Million Reddit Questions, One Million Reddit Jokes, One Million Reddit Confessions, Legal Advice Reddit, Five Years of AAPL on Reddit, FigLang 2020 Reddit Dataset, CodeSwitch-Reddit, lmqg/qg_squadshifts, Reddit
This dataset is used in 1 benchmark:
Recent papers with results on this dataset: