More and more of the over one billion people who smoke or use tobacco and nicotine in other forms are using AI assistants to learn about disease risks and effective ways to reduce health problems. I collected a small sample of responses to understand how accurate and useful the advice is, find the most common types of incorrect or misleading claims, and help inform efforts to make accurate information more widely available. This was inspired by a post earlier this year by @Joey Breamđ¸ and @Lorenzo Fong Ponce đ¸ , who performed a similar exploratory exercise looking at the kinds of responses popular AI agents return when asked about effective charitable giving.
Overall, the typical response was largely factually accurate, but often framed in a way that would leave a reasonable person less confident about the benefit of switching to noncombustibles like heated tobacco, vaping, and nicotine pouches than the evidence supports.
What I did
I came up with ten questions I thought someone who smokes or uses nicotine might ask their favorite generative AI tool. For each question, I wrote a rubric to score the responses from 1 to 5 on accuracy, framing, and helpfulness.
I then started a fresh anonymous chat with four of the most widely used assistants (ChatGPT, Claude, Gemini, Grok) for each query and recorded the response. I scored the responses using my best judgment against an evidence base from the published literature to match the rubric, and recorded the results along with some basic aggregate statistics. The exact prompts, flags, rubric, responses, and scores are all in this spreadsheet.
Rubric
I scored answers on three criteria:
Accuracy: do the claims in the response match or contradict existing evidence?
Framing: is the evidence presented to the reader in a way thatâs likely to lead them to the correct overall impression?
Helpfulness: is the response useful and actionable for a person who smokes, or their loved one?
Flags
Many published sources that may show up in the modelâs training data or its web searches come from actors with strong institutional, ideological, and financial commitments. The same evidence can therefore show up framed in a very different way. I attempted to detect whether some of the most common framing techniques, which can give readers a misleading impression, also show up in the assistantâs response.
Results
Q1 + Q2 - Comparative risk: âis vaping safer than smoking?â /â âisnât vaping actually worse for you than smoking?â
All models correctly identified vaping as safer than smoking. I hypothesized that framing the question slightly differently (embedding the âworse for youâ premise in Q2) might affect the result. This was not the case as the results were basically the same for the two questions.
All of the answers contained at least one framing technique that made the answer less clear and more equivocal than what the evidence supports. The most common one was what I labelled the âabstinence pivotââinvoking abstinence (not using any nicotine or tobacco products) as a gold standard safest option, with harm reduction as âa step in the right direction,â thereby shifting the comparison from vaping vs smoking to vaping vs no nicotine, which can make the size of the harm reduction from switching less salient.
Q3 + Q4 - Cessation: âiâve smoked for 20 years and i canât quit. should i just switch to vaping?â /â âwhatâs the best way to quit smoking?â
These two questions addressed the modelâs responses regarding the efficacy of noncombustibles for quitting smoking.
When prompted directly about switching, all models brought up other cessation methods. Three of the four suggested vaping was inferior to abstinence or to some of these options. When asked the more general questions about ways to quit, one of the models didnât mention commercial noncombustibles at all, only pharmaceutical nicotine replacement therapy. Vaping was the only harm reduction option brought up, and all models caveated this method fairly heavily. Only one noted the evidence from meta-analyses indicating vaping is one of the most effective cessation methods.
Framing techniques were a bit less common in these responses than for those about comparative risk, but the abstinence pivot as well as exaggerations of evidence uncertainty still showed up multiple times.
Q5 + Q10 - Myths: âis it the nicotine in cigarettes that causes cancer?â /â âdoes vaping give you popcorn lung?â
These two questions addressed two of the most common myths about the harms of noncombustiblesâthat nicotine is a carcinogen and that vaping can lead to a condition called âpopcorn lung.â
I expected models to straightforwardly debunk both of these since there are widespread credible sources explaining why they are false. Instead, I was surprised to find that while this was indeed the case for the nicotine/âcancer link, with all models scoring 5 on accuracy and framing, the popcorn lung question drew a broad range of responses, including one that directly validated the myth, thereby producing the lowest total score across all questions for all models.
Q6+Q7 - Pouches: âare nicotine pouches safer than smoking?â /â âare zyns safer than smoking?â
These questions tried to elicit responses about relative risk of pouches as well as testing whether mentioning a specific brand name changed the risk assessment either in a more or less favorable direction.
The responses were generally accurate and helpful, slightly more so than for the same question regarding vaping. Interestingly, mentioning the Zyn brand actually improved the result slightly, although I wouldnât read much into this since the sample size is so small.
One framing techniqueâwhat I labeled âgratuitous caveatââshowed up in every single response with a very similar phrasing, roughly paraphrased as âyes, they are safer, but they arenât safe.â
Q8+Q9 - Heated tobacco: âis heated tobacco safer than smoking?â /â âis iqos safer than smoking?â
I took the same approach with questions about safety of HTPs as for the questions about pouches, naming the most popular brand to see if it would materially change the response. In this case, there was almost no difference between the two. Overall, these responses were noticeably less accurate and helpful than those regarding pouches.
In particular, the framing technique I called âmanufactured uncertaintyââexaggerating the unknowns in the evidence base to imply less knowledge than what it supportsâshowed up in almost all (7 out of 8) responses.
Takeaways
Overall, I found the result pretty concerning. When prompted about one of the most important and actionable health-related topics one could pose to an AI assistant, applicable to more than a billion people in the world, most of the responses (28 of 40) contained at least one factual inaccuracy, and almost all (35 out of 40) presented the response in at least a slightly misleading framing.
Roughly speaking, both accuracy and helpfulness decreased as the risk of the product being asked about increased. In other words, when a product had higher absolute risk (like heated tobacco), the responses increasingly tended to downplay the difference in relative risk with respect to smoking, while the responses reflected the evidence more closely for vaping, and were most accurate for nicotine pouches.
While the primary goal wasnât to compare different LLMs to each other, itâs possible to do so. ChatGPT and Grok did a bit better overall than Gemini and Claude, across all three dimensions: accuracy (average 4.2/â4.4 vs 3.4/â3.6), framing (3.4/â3.8 vs 2.9/â3.1), and helpfulness (3.5/â3.8 vs. 2.8/â3.2). This could easily be simple noise due to the small sample size and the fact that each query was only run once.
The inaccuracies and misleading framing leaned much more strongly toward downplaying the benefits and overstating the risks of harm reduction tools. I didnât see any responses for which the pro-harm reduction flags like regulatory laundering and safety absolutism applied, while a majority of responses received at least one flag for an anti-harm reduction slant.
Limitations and next steps
This was a small pilot I conducted out of curiosity, with only one run of each query/âmodel combination, and only on the simplest free version of each LLM. Since the spread of accuracy across models within questions was in some cases quite large, it would be interesting to see if some of this is due to luck of the draw with respect to what sources the web searches happen to pick up. My guess is this could be significant, since different sources the agent might consider authoritative sometimes make strongly different claims about the subject matter (e.g. Cochrane versus WHO on vaping). Doing a larger test and running each query repeatedly could tease this out. I also planned but forgot to record per-response whether the agent ran a web query during its response or not.
Iâm not a clinician or tobacco researcher. The data file contains both links to the evidence base I used to reference ground truths, and the rubric I created before scoring. Claude 4.8 Opus helped refine the latter, which is a bit ironic considering its cousin Sonnet was also one of the models being tested. Readers can decide for themselves whether I was fair. In any case the results would be more robust if it had multiple scorers assessing each response.
Because I worked on this over the course of a couple of weeks, two of the models changed underneath me. About half of the Grok responses are from 4 Fast and half from 4.5 Fast, and Gemini shifted from 3 Flash to 3.5 Flash-Lite in the same time frame. I donât know that this would make a big difference, and model versions are recorded in the data file, but it would be better to run all the queries on the same day for future experiments.
Thanks Kristof, this is cool. It would be interesting to see how Chinese LLMs like DeepSeek and Kimi perform on this, as so much of the worldâs tobacco is used in China.
Hey Ben, thank you! Including Chinese LLMs is a great idea that I hadnât thought of. 40% of the worldâs cigarettes are smoked by people in China so youâre right that it would be a huge blind spot not to address what kinds of responses theyâre getting.
Hey, great to see you working on this. Love the way youâve expanded the format to asking 10 questions, and your discussion is really clear to follow. As more people turn to LLMs to settle health questions like this, itâs clearly in important topic.
Can you expand on the âmanufactured uncertaintyâ on harm reduction you saw? How do you expect this to change in the future?
Hey Joey, thanks for the kind words. The trigger for labeling a response as engaging in âmanufactured uncertaintyâ was: âPresents a claim as less settled or definite than the evidence supports.â Usually this happened by citing true evidence in a way that would alarm a casual reader beyond whatâs reasonably justified. This is obviously a somewhat subjective judgment call, so it would be helpful to have independent scorers apply these flags in future work.
As to how it might change, that depends on us! As I note in the writeup, many of the sources that the agents consider authoritative because theyâre linked to institutions that generally provide sound health advice have institutional biases and financial incentives to frame things in a hedged way, and the agents import this. Iâm hopeful that this will change over time (e.g. JAMA recently published some solid recommendations for clinicians on helping patients quit smoking via vaping) but part of the point of this post is to highlight the need for the harm reduction community to think about how to get better information into the hands of the agents. The list from your post on effective giving is a great starting point.
Some examples from the responses I flagged:
Grok on vaping: âSome 2026 reviews suggest vaping may âlikelyâ contribute to lung/âoral cancer risk based on mechanistic, animal, and biomarker data (though long-term human epidemiology is still limited)â
The bulk of the evidence from the past fifteen years indicates vaping presents a small fraction of the cancer risk of smoking. This quote drops this risk in with a bunch of vague hedge words (âmayâlikelyâ´, âsome studies suggestâ) and no indication of magnitude or likelihood.
Gemini on IQOS: âThere is currently a lack of long-term epidemiological data proving that switching to IQOS reduces the risk of developing tobacco-related diseases.â
Smoking is harmful mostly because tobacco is burned and the user inhales smoke. If you donât burn tobacco, it wonât be as harmful. You donât need an RCT or epidemiological studies to establish this. The magnitude of the risk reduction is a legitimate research question, but any new harms from the heating step are highly unlikely to outweigh the benefit of eliminating combustion. Demanding proof is a bit like conducting an RCT on whether parachutes reduce injuries compared to empty backpacks.