- Future//Proof
- Posts
- π One in three AI citations doesn't say what the AI claims it says
π One in three AI citations doesn't say what the AI claims it says
π§Ύ Plus: Mistral trains on free-tier prompts, and a 17x cost gap.
Welcome to the Future//Proof π
π Hello , the AI Enthusiast.
In this weekβs edition, we brought AI updates backed by high-quality research and data to give you deeper insights. You'll find the Top AI Breakthrough of the Week, a featured AI tool with a mini-tutorial, learning resources to help you master these tools, the top 3 AI news stories, and more.
Our goal is to help you improve your knowledge and stay ahead in the rapidly evolving AI landscape. You can submit your questions, queries, thoughts, opinions or anything regarding AI as a reply to this email and we'll feature and address them in our next newsletter.
π Now Letβs dive in and explore the new AI Insights together!
β Read Time: 6m:41s
Granola Runs Revenue On Attio
"When I think of revenue, I think of Attio." - Shreman Shrestha, Head of Business at Granola
Here's what that adds up to:
Zero missed leads and 10x faster access to customer context
Lead triage 83% faster
Five hours saved per week with automated updates

An in-depth look at a major AI development, its industry impact, how it could affect your career, and a bold future prediction.

Two audits just took AI citations apart, and the results are ugly
Two independent research groups published audits of AI answer engines this week, and both landed on the same uncomfortable conclusion from different angles. Haus Research put 310 factual questions about tech companies to Perplexity's sonar and sonar-pro models at temperature zero, then did the thing almost nobody does: it fetched all 2,915 cited URLs and checked them. 34.7% of citations pointed to a page that either would not open or contained none of the numbers being attributed to it. Per individual claim the failure rate was 14.4% across 872 claims. The breakdown matters more than the headline, because only 1.3% were genuinely dead links, while 16.1% sat behind paywalls and another 16.1% were live, loading, perfectly real pages that simply did not contain the figure. Identifying a company's CEO was the weakest category at 44.3% accuracy, and a quarter of all cited URLs had never been archived by the Wayback Machine, meaning they cannot be checked later. The second audit, from researcher Trellner, asked why the sources are so odd in the first place. Across 380 software queries it collected 7,534 citations and found that nearly 60% came from sites outside the global top 100,000. Three domains in particular, wifitalents.com, worldmetrics.org and gitnux.org, together host 215,128 machine-generated buying guides, share identical page templates, and were all registered through the same registrar between December 2023 and May 2024. That is not accidental low quality. That is a content operation built to be cited by AI, not read by humans. Both audits are researcher-run rather than peer-reviewed and both focus on Perplexity's Sonar models specifically, so treat the exact percentages as directional. The pattern, however, is not in dispute.
Potential Impact
The uncomfortable part is not that AI makes mistakes, which everyone already knows, but that the citation was supposed to be the fix, and this is the first serious evidence that the fix does not hold. A footnote next to a sentence creates a powerful feeling of verification, and roughly a third of the time that feeling is unearned. Worse, the failure mode is almost invisible: a link that loads a real, credible-looking page is far more dangerous than a broken one, because nobody clicks a working link to check whether the number is actually on the page. Layer the Trellner finding on top and the problem compounds, because if a coordinated network can manufacture 215,000 pages specifically to be harvested by answer engines, then the citations are not just unreliable, they are targetable. Anyone who understands that AI engines reward machine-readable, statistic-dense, well-templated content can manufacture the appearance of authority at industrial scale, and the audits suggest some already have. For anyone whose work involves stating numbers to clients, investors, regulators or customers, that is a live professional risk, not a philosophical one.
Implications for People/Careers
The practical rule for every working professional is simple and slightly boring: an AI citation is a lead, not a source. If a number is going into a proposal, a report, a pitch deck, a board paper or anything with your name on it, open the link and find the figure on the page with your own eyes. It takes fifteen seconds, and this research says roughly a third of the time you will be glad you did. Watch particularly for numbers that are current-state facts, since the CEO identification result at 44.3% accuracy suggests answer engines are weakest exactly where the web is most out of date. For business owners there is a second, less obvious implication, and it may be the more valuable one: AI answer engines have quietly become a discovery channel, and right now that channel is being won by content farms rather than by real businesses. If a potential customer asks an AI which provider they should use in your category, something is being said about you or, more likely, said instead of you. Almost nobody has checked what that is. The competitive gap here is not clever prompting, it is being a primary source: publishing your own real numbers, clear pricing, dated case studies and structured, factual pages that an engine can lift and cite correctly.
Our Future//Take
Our prediction is that the next twelve months bring a visible tightening of source quality in AI answer engines, in the same way search engines eventually came for link farms, and that the labs will start treating source reputation as a ranking problem rather than a retrieval problem. Expect citation-verification to become a product feature, where the engine checks that the claimed figure is actually on the cited page before showing it to you, because the audits published this week make it embarrassing not to. In the meantime, two things are worth internalising. The first is that verification is now a billable professional skill, and the people who quietly build the habit will avoid a category of career-damaging mistake that is going to become common. The second is that being genuinely, verifiably correct on your own website is turning into a marketing asset rather than a compliance chore. The businesses that publish real, dated, specific numbers are the ones the next generation of answer engines will be able to trust, and trust is about to become the scarce commodity in a channel currently flooded with 215,000 fake buying guides. Hereβs your βΉ25,000 AI Gift for FREE π

Quick summaries of this week's top AI news, their relevance to your career, and our expert opinions.
Adobe launched Adobe for Slack on September 2, an MCP-based bot that exposes more than 70 Adobe tools including Firefly, Photoshop, Express, Premiere, Acrobat, InDesign, Illustrator, Stock and Lightroom from inside a Slack conversation. You describe what you want, the bot routes to the appropriate Adobe tool, and it uses the surrounding thread context to generate PDFs, images and video without anyone opening a desktop app. It is available to Slack Business+ and Enterprise+ teams.
Why It Matters to You
This is the clearest sign yet that MCP has stopped being a developer acronym and started being a distribution strategy. We flagged this protocol two editions ago as the plumbing that lets any AI call any tool, and here is a company with a large creative software estate using it to move its products into the place where work is actually discussed. The practical effect for a marketing or content team is the removal of a handoff, since the request, the asset and the feedback now live in one thread instead of crossing a brief, a designer's queue and a review link.
Our Take
Two thoughts. First, the constraint is real, and Business+ or Enterprise+ Slack plus Adobe licensing puts this out of reach for a lot of smaller teams, so read it as a signal about direction rather than a tool you will deploy tomorrow. Second, and more useful, ask which of your own tools your team currently leaves the conversation to use. That switch is where hours quietly disappear, and MCP-based connectors are increasingly available for CRMs, ticketing systems and project tools. The winning move is not adopting Adobe's version of this, it is noticing the pattern and applying it to whichever tool your team opens twenty times a day.
Mistral updated its help centre this week to state explicitly that free-tier Vibe conversations are used to train its models unless a user manually opts out in the admin panel. Vibe Enterprise, Mistral Studio and API traffic are opted out by default. The detail that caught developers' attention on Hacker News is that the Vibe and API toggles are independent, so opting out of one does not cover the other, and the clarified default sits awkwardly against earlier documentation wording.
Why It Matters to You
Almost every team has someone using a free AI tier for something they should not be. Draft contracts, client lists, unreleased pricing, internal financials, candidate CVs. The Mistral clarification is not a scandal, since this default is common across the industry, but it is a useful prompt to go and check, because the risk is not the vendor, it is the assumption. Most people genuinely believe their prompts are private, and on free consumer tiers they usually are not by default.
Our Take
Run a fifteen-minute audit this week. List every AI tool anyone on your team touches, note which tier each is on, and find the training toggle for each one. Then write down the actual rule for your business, which for most is simple: anything containing client data, financials or unreleased information goes only into a paid or enterprise tier with training disabled, and everything else can live wherever. Bear in mind that per-product toggles are becoming the norm rather than one account-level switch, so check each surface separately. This is the cheapest risk reduction available to you right now.
An eval published this week ran 360 identical software-engineering trials on a single model across 9 different tools, holding the infrastructure constant. The result: cost per successful task ranged from $1.05 to $18.34, roughly a 17x spread on the same underlying model, and the expensive end was not proportionally better, delivering 63.3% success against 53.3% for the cheapest. Separately, an analysis published the same day argued that a low-cost model at roughly $0.02 per task handles about 90% of real workloads about as well as a premium model at $3.69, a gap of more than 100x. Google also shipped Gemini 3.8 Flash this week at $0.75 per million input tokens and $3.75 per million output tokens through December 31.
Why It Matters to You
Most AI budget conversations focus entirely on which model to buy. This data says the wrapper around the model may matter more to your bill than the model itself. The tool decides how many times it calls the model, how much context it resends, how often it retries, and how much it wastes on approaches that fail. Two teams using the identical model through different software can have wildly different invoices for the same output.
Our Take
The practical version for a non-technical owner is this: when you evaluate an AI vendor, stop asking which model is under the hood and start asking what a completed unit of work costs. Cost per resolved ticket, per qualified lead, per generated report. That single number captures the model, the tool and the waste all at once, and it is the only figure that connects to your P&L. Ask any vendor to quote it, and be sceptical of anyone who cannot. AI Mastery for FREE (Sign up Now).

Discover a comprehensive guide to an AI tool, exploring its features, practical use cases, and learning resources to help you master it.

π Perplexity
Yes, we just spent the Big Picture pulling apart research into Perplexity's citations. That is precisely why it is this week's tool. It remains one of the most useful research instruments available to a working professional, and this week gave us an unusually clear map of exactly where it breaks. A tool you understand the failure modes of is safer than one you trust blindly.
Perplexity is an answer engine rather than a chatbot. You ask a question, it searches the live web, and it returns a written answer with numbered citations attached to each claim. For competitive research, market sizing, vendor comparisons and catching up on a topic quickly, it is faster than anything else. The discipline this week's research demands is that you treat the answer as a well-organised starting point rather than a finished one.
β Top Features
Live web answers with citations. Every claim carries a numbered source you can open, which is what makes structured verification possible at all.
Spaces. Persistent research collections per topic, with their own uploaded files and instructions. Create one per competitor, client or project instead of losing everything in chat history.
Deep Research. Runs an extended multi-search investigation and returns a longer report. Best used for landscape questions rather than single facts.
Source filtering. Restrict a search to academic sources, or to social discussion, depending on whether you want rigour or real opinion.
Custom instructions. Tell it your role and standards once. Instructing it to say when it is uncertain rather than guessing measurably improves what comes back.
Comet browser. Perplexity's free Chromium-based browser with the assistant built in, able to read and summarise the page you are on.
Resources for Learning
Official Help Centre: perplexity.ai/help-center for Spaces, source filtering and current plan limits.
The Citation Audit: hausresearch.com for the full methodology behind this week's findings, which is worth skimming before you rely on any answer engine.

A curated list of noteworthy AI tools and their key details to help you stay ahead in your field.

Meta's first real-time speech model, launched this week, handles streaming transcription with speaker separation and processes audio in 80ms chunks, holding up across recordings over an hour long with 20+ speakers. It was trained across 70+ languages with production support for 25+, and Meta reports a 3.1% word error rate on streaming with 17.5% diarization error. It ships through the Meta Model API and Meta AI for Mac. The differentiator is multi-speaker accuracy in real time, which is the specific thing that breaks most transcription tools the moment a meeting gets crowded or someone interrupts.

Google's newest workhorse model shipped this week at $0.75 per million input tokens and $3.75 per million output tokens through December 31, pitched partly as a fix for verbose output, which quietly matters because you pay for every unnecessary word. A security-focused variant, Gemini 3.8 Flash Cyber, launched alongside it but is restricted to vetted defenders. The differentiator is cost per useful answer rather than raw capability, which is exactly the axis this week's cost research says most buyers are ignoring. Worth testing against whatever you currently run for high-volume, low-complexity tasks.

An AI platform that conducts customer interviews in 15 languages and analyses the transcripts for market insight, which raised a $50M Series A led by DST Global this week. It sits in the space traditional survey and interview vendors used to own, replacing weeks of scheduling and manual coding with conversations that run on demand. The differentiator is depth at survey scale: open-ended conversation rather than multiple choice, across a sample size you could never interview by hand. If you have been putting off proper customer research because of the time cost, this category is worth a look.

A free browser tool Anthropic launched this week that reads embedded C2PA credentials and tells you whether a file was created or edited with Claude. It accepts images, video and audio up to 100MB. Crucially, Anthropic is upfront that a missing signal does not prove AI was not used, since credentials can be stripped or simply absent. The differentiator is honest provenance: it is one of the first mainstream tools to tell you what it cannot prove, which given everything else in this issue is a refreshing standard to hold tools to.

A quick poll to help you recollect and engage with key points from the newsletter.
In this week's citation audit, roughly 35% of AI citations failed verification. What was the single biggest cause? |

Share your feedback on today's edition to help us improve and better meet your needs.
How was Todayβs Edition? |
Share our Newsletter β©
Enjoying our insights on the latest AI breakthroughs? Donβt keep it to yourself! Share this newsletter with friends and colleagues who are passionate about technology and AI innovation.
If you havenβt subscribed yet, make sure to subscribe here to stay updated with cutting-edge AI news, tools, and tutorials delivered straight to your inbox!
Ask Us Anything AI β
Got questions? We've got answers!
Submit your questions, queries, thoughts, opinions or anything regarding AI and we'll feature and address them in our next newsletter. Your curiosity drives our content!
π Reply to this email with your questions, and we'll answer them in our next edition!π

