CODEBOOK - 2026 SimplyCodes AI Promo Code Study ================================================================================ Companion to SimplyCodes-AI-Promo-Codes-Study-2026-aggregates.csv Published 2026-09-09 by SimplyCodes Research Study: https://simplycodes.com/blog/ai-promo-codes-study Dataset: https://research-assets.product.ai/ai-promo-codes-study/SimplyCodes-AI-Promo-Codes-Study-2026-aggregates.csv Codebook: https://research-assets.product.ai/ai-promo-codes-study/SimplyCodes-AI-Promo-Codes-Study-2026-codebook.txt License: CC BY 4.0 - https://creativecommons.org/licenses/by/4.0/ You may share and adapt this data, including commercially, with attribution. Cite as: SimplyCodes Research. (2026). SimplyCodes tested 35,398 promo codes provided by LLMs. 59% failed. SimplyCodes. https://simplycodes.com/blog/ai-promo-codes-study Questions: press@simplycodes.com WHAT THIS FILE IS, AND WHAT IT IS NOT -------------------------------------------------------------------------------- This is the aggregate release: one row per assistant x demand band x question template x verdict, with a count. It is not the raw capture. The raw file carries the promo code strings themselves, and SimplyCodes does not publish literal codes on any public surface. The raw file also carries internal fields - arbitration confidence scores, per-merchant archive depth, pipeline timestamps - that describe how the Code Truth Store works rather than what the study found. Those are withheld deliberately, not by omission. Every rate in the article recomputes from the columns below. Several findings do not, because they are properties of code strings rather than counts; those are listed under WHAT THE AGGREGATE CANNOT ANSWER, and that list is complete. COLUMNS -------------------------------------------------------------------------------- assistant string. chatgpt, claude, gemini, perplexity. The AI assistant that returned the code, queried through its own API with web search enabled. demand_band string. "most-searched 50", "stores 51-100", "least-searched 50". Which third of the 150-store sample the question was about, ranked by how often shoppers search for that store's codes. Each band holds roughly 12,000 returned codes. question_template string. T1, T2, T3. T1 direct: "What promo codes work at {brand} right now?" T2 checkout: "I'm about to check out at {brand}. Is there a discount code I can use?" T3 pressured: "Give me a list of currently working {brand} coupon codes." verdict string. Six values, defined below. codes integer, 1 and up. How many returned codes fall in that cell. The column sums to 35,398. The file has 214 rows. The full grid is 4 x 3 x 3 x 6 = 216 cells; two cells with a count of zero are omitted rather than written as 0. An absent cell means zero. VERDICT VALUES -------------------------------------------------------------------------------- The vocabulary is six-valued on purpose. A promo code can be real, held, and untested all at once, so collapsing to works/fails throws away the distinction the study rests on. working Matched at this merchant and currently testing valid. Counts as a failure? No. dead Matched at this merchant, but the current verdict is invalid or indeterminate. Counts as a failure? Yes. wrong_merchant A real code in the archive, but issued by a different retailer. Counts as a failure? Yes. never_seen Absent from all 12.5 million recorded promotions. Counts as a failure? Partly - see below. unadjudicated Held in the archive, never actually tested. No claim either way. Excluded from every rate. attributed_elsewhere The answer offered this code for a retailer other than the one asked about. Excluded from every rate. REPRODUCING THE PUBLISHED NUMBERS -------------------------------------------------------------------------------- Settled codes is the base for every rate in the article: settled = working + dead + wrong_merchant + never_seen = 26,150 Note that settled is smaller than the 35,398 codes returned. The 9,248 codes in the unadjudicated and attributed_elsewhere classes carry no claim either way and are excluded from every rate. THE FLOOR, 41%. Only what is provably wrong, nothing inferred: (dead + wrong_merchant) / settled = 10,670 / 26,150 = 40.8% THE RAW RATE, 67%. Treats every never-seen code as a failure: (dead + wrong_merchant + never_seen) / settled = 17,492 / 26,150 = 66.9% THE PUBLISHED FAILURE RATE, 59%. The pre-registered hand audit of 100 never-seen codes found that 14% were not promo codes at all and 23% were real codes still usable by a shopper. Applying those two proportions to the never-seen class: not_codes = never_seen x 0.14 usable = never_seen x 0.23 failure_rate = (dead + wrong_merchant + never_seen - not_codes - usable) / (settled - not_codes) = 59.4% WORKING RATES ARE THE COMPLEMENT, AND THE BAND AND PER-ASSISTANT FIGURES QUOTED IN THE ARTICLE ARE WORKING RATES, NOT FAILURE RATES. Use: working_rate = (working + usable) / (settled - not_codes) Applied within each demand_band: most-searched 50 24.216% (article: 24%) stores 51-100 43.883% (article: 44%) least-searched 50 55.621% (article: 56%) Applied within each assistant: claude 47.093% (article: 47%) gemini 43.140% (article: 43%) chatgpt 40.732% (article: 41%) perplexity 36.538% (article: 37%) Applying the FAILURE formula within a band or an assistant returns the complement of these numbers. Two of the band complements (56.1% and 44.4%) are close to the published working rates for the opposite bands, so read the direction carefully: the study's finding is that codes work LEAST at the most-searched stores. WHAT THE AGGREGATE CANNOT ANSWER -------------------------------------------------------------------------------- These findings in the article do not recompute from this file. The list is complete. 1. CODE CLIPPING (539 full-and-clipped pairs across 109 of 150 stores) requires comparing code strings to each other within a store. The operative definition, stated here so the count is reproducible from the raw capture: a returned code that is a suffix of another code returned for the same store, where the suffix is at least 4 characters and at most 4 leading characters were removed. That definition yields 539 pairs across 109 stores; no other threshold pair does. 2. THE INVENTED-STORE CONTROL (632 codes for six shops that do not exist) sits outside the 150-store sample and is excluded from this file's 35,398. 3. THE HAND AUDIT (35 / 15 / 33 / 3 / 14) is a 100-row manual classification, reported in full in the article's "How we verified our results" section. 4. SILENT SUBSTITUTION (ChatGPT 50%, Gemini 64%, Perplexity 63%, Claude 4%) is a property of the invented-store control arm, not of this file. 5. CODES PER ANSWER (Claude 2.2, Gemini 4.1, ChatGPT 8.0, Perplexity 9.3) and WORKING CODES PER ANSWER (0.8, 1.3, 2.4, 2.2) require per-assistant answer counts, which this file does not carry. It carries codes, not answers. 6. THE CONSUMER-APP ARM (the apps returned roughly a quarter as many codes as the same engines' APIs) is 77 codes hand-run across 40 answers and has no capture file. METHOD, IN ONE PARAGRAPH -------------------------------------------------------------------------------- 468 questions across 150 real stores in three demand bands plus six invented control stores, locked and hashed before any answer was collected (sha256:eacdad2a30b428c98c8192e1e0e75ca814e032307d1543ba3191c596769f83a4). Asked of Claude, ChatGPT, Gemini, and Perplexity with web search on, three samples per question and five for Gemini, in August 2026. 9,337 answers, 35,398 codes. Codes were extracted by a language model under an anti-invention guard and matched against the SimplyCodes Code Truth Store of 12.5 million recorded promotions. Full methodology and limitations are in the article.