Technology
Researchers Use Reddit to Study Cannabis and Pain

When people talk about cannabis and pain relief, the conversation often focuses on clinical trials, patient surveys, or personal testimonials. But what if researchers could learn from thousands of real-world conversations already happening online?
A 2026 study1 published in Data in Brief takes a unique approach to understanding cannabis-related discussions. Rather than testing whether cannabis reduces pain, researchers created a specialized dataset using publicly available Reddit posts to help train artificial intelligence tools that can better analyze how people talk about cannabis in the context of pain management.
The project doesn’t answer whether cannabis works for pain. Instead, it provides a valuable research resource that could help scientists better understand patient experiences, identify emerging trends, and improve the way health-related discussions are analyzed online.
Why Cannabis Discussions Matter Online
Millions of people turn to social media platforms to share their health experiences. Reddit, in particular, has become a popular destination for discussions about chronic pain, autoimmune conditions, medications, and alternative treatment options. These conversations can offer researchers a glimpse into patient perspectives that may not always appear in traditional medical studies. People often discuss symptoms, treatment decisions, side effects, and personal experiences in their own words.
However, analyzing large volumes of online conversations is challenging. Human researchers can only review so many posts, which is why artificial intelligence and machine-learning systems are increasingly being used to help process health-related data. Before AI systems can accurately interpret these conversations, they need high-quality training data. That’s where the cannabis Reddit dataset comes in.
Building a Cannabis Reddit Dataset
The research team developed a manually annotated dataset focused on cannabis-related discussions within Reddit communities associated with autoimmune rheumatic diseases. To create the dataset, researchers searched publicly available posts using a structured collection of cannabis-related terms. After identifying relevant discussions, they extracted specific text segments and carefully reviewed them. The final dataset included 479 post-aspect pairs, meaning individual pieces of text were linked to particular cannabis-related topics or aspects being discussed.
What makes this dataset especially useful is that it was manually annotated by human reviewers rather than relying solely on automated tools. Human annotation helps create a benchmark that future AI systems can learn from and be evaluated against.
Researchers then categorized the content according to sentiment, allowing future machine-learning models to learn how people express positive, negative, or neutral views about cannabis-related experiences.
What Is Aspect-Based Sentiment Analysis?
One of the most interesting parts of the project is its focus on something called aspect-based sentiment analysis. Traditional sentiment analysis looks at the overall tone of a statement. For example, a post might be classified as generally positive, negative, or neutral. Aspect-based sentiment analysis goes a step further. Instead of evaluating an entire post as a single unit, it identifies specific topics within the discussion and determines how the writer feels about each one individually.
Imagine someone writes that cannabis helped them sleep better, but caused unwanted side effects. A traditional system might struggle to categorize the post because it contains both positive and negative opinions. Aspect-based sentiment analysis can separate those viewpoints and recognize that the person expressed positive feelings about sleep improvement while expressing negative feelings about side effects. This more detailed approach may help researchers better understand the complexity of health-related conversations.
What Researchers Found
The study’s primary goal was not to evaluate cannabis as a treatment but to create a reliable dataset for future research. To assess the quality of their annotations, the researchers measured agreement among reviewers using a statistical method known as Krippendorff’s alpha. This helps determine how consistently different people interpreted the same content.
The results showed moderate agreement among annotators for both traditional sentiment analysis and aspect-based sentiment analysis, suggesting that the dataset provides a useful foundation for future machine-learning development. More importantly, the project demonstrated that cannabis-related discussions can be systematically identified, organized, and labeled in ways that support advanced natural language processing research.
The dataset is now publicly available, allowing other researchers to use it for training, benchmarking, and evaluating AI models designed to analyze health-related social media discussions.
Why This Dataset Is Important
At first glance, a dataset paper may not sound as exciting as a clinical trial. Yet resources like this often play a critical role in advancing scientific research. Artificial intelligence systems are only as good as the data used to train them. Without carefully curated datasets, researchers risk developing models that misunderstand patient experiences or fail to recognize important patterns in health discussions.
This dataset may help improve future tools used in:
- Public Health Informatics
- Digital Epidemiology
- Medical Research
- Natural Language Processing
- Social Media Health Monitoring
As AI becomes increasingly integrated into healthcare research, high-quality datasets can help scientists better analyze large-scale patient conversations while reducing the amount of manual review required.
What This Study Does Not Show
Because cannabis research often attracts significant public attention, it’s important to understand what this study was designed to accomplish, and what it was not.
The dataset does not demonstrate that cannabis reduces pain. It does not prove that cannabis is effective for autoimmune diseases, arthritis, rheumatic conditions, or any other medical condition. Nor does it evaluate treatment outcomes, compare cannabis products, or measure symptom improvement. Instead, the study focuses entirely on creating a structured collection of annotated Reddit posts that future researchers can use when developing and testing AI systems. The project is about understanding language and patient discussions, not evaluating medical efficacy.
The Future of AI and Cannabis Research
As healthcare conversations increasingly move online, researchers are looking for new ways to learn from the experiences people voluntarily share on digital platforms. Projects like this highlight how artificial intelligence and public health research are beginning to intersect. Rather than replacing traditional clinical studies, these tools may eventually complement them by helping researchers identify trends, concerns, and emerging topics worthy of further investigation.
The real value of this dataset lies in its potential to support future discoveries. Giving researchers a reliable resource for training and evaluating machine-learning models helps build the foundation for more sophisticated analyses of patient-reported experiences.
While the study doesn’t tell us whether cannabis works for pain, it does reveal something equally important: understanding how people talk about cannabis may be a crucial step toward understanding how they experience it. As AI tools become more advanced, datasets like this could help researchers turn vast amounts of online conversation into meaningful scientific insights, one post at a time.
References:
1. Tricia Park, Sahithi Lakamana, Aishwarya Alagappan, Catherine Diop, Ege Gursel, Yuting Guo, Titilola Falasinnu, Abeed Sarker, Selen Bozkurt, A Cannabis Use Reddit Dataset for Aspect-Based Sentiment Analysis, Data in Brief, 2026, 112907, ISSN 2352-3409, https://doi.org/10.1016/j.dib.2026.112907












