- Future//Proof
- Posts
- π¨ An AI agent just ran a data breach start to finish, and a regulator has the paperwork
π¨ An AI agent just ran a data breach start to finish, and a regulator has the paperwork
ποΈ Gemini tops voice, and your AI limits just shrank.
Welcome to the Future//Proof π
π Hello , the AI Enthusiast.
In this weekβs edition, we brought AI updates backed by high-quality research and data to give you deeper insights. You'll find the Top AI Breakthrough of the Week, a featured AI tool with a mini-tutorial, learning resources to help you master these tools, the top 3 AI news stories, and more.
Our goal is to help you improve your knowledge and stay ahead in the rapidly evolving AI landscape. You can submit your questions, queries, thoughts, opinions or anything regarding AI as a reply to this email and we'll feature and address them in our next newsletter.
π Now Letβs dive in and explore the new AI Insights together!
β Read Time: 7m:04s
Some teams never seem to stop moving. They're on Attio, the agentic CRM.
Every customer signal is captured in one shared context layer, always current and compounding. Agents and workflows build pipeline, chase every buying signal, and move deals forward, an always-on revenue engine running alongside your team.
With Attio, youβll get:
Leads automatically prioritised and routed to the right rep
Expansion and risk signals caught the moment they land
Follow-ups written in your voice, already there when you arrive
Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?

An in-depth look at a major AI development, its industry impact, how it could affect your career, and a bold future prediction.

Spain just logged the first data breach carried out end to end by an AI agent
Two weeks ago we wrote that a credential-harvesting campaign now takes six hours to build rather than six weeks. This week a European regulator filed the paperwork on what that actually looks like. On September 15 the Spanish Data Protection Agency disclosed what it describes as the first personal-data breach executed end to end by an AI agent operating outside a laboratory. A third party pointed a well-known large language model at a Spanish organisation, and from there the agent chained reconnaissance, login, application probing, data modification and access to invoices without human steering. Nobody sat at a keyboard directing each step. The AEPD was candid that it cannot yet say where the agent came from, floating three plausible origins: a jailbroken guardrail on a commercial model, a sandbox escape from a testing environment, or a penetration tester's custom model built on a popular LLM. Hold that uncertainty, because this is a breach notification rather than a completed forensic report. What is not uncertain is the category shift. Agentic attacks have moved from conference demonstrations into a regulator's formal filings, which is the point at which insurers, auditors and courts start paying attention. A second disclosure this week made the same argument from the defensive side, when security firm Strix published details of an agent that found an exposed container registry belonging to an AI infrastructure company, pulled an image, and located a GitHub token preserved in Docker build history since March 2023 that still carried admin and push access to the company's main product repository, its deployment repository and its software distribution channel. The token was rotated within hours of the July report and the company consented to publication, so nothing was lost. The lesson is what an agent found in a place nobody had thought to look at for three years.
Potential Impact
The useful way to read these two stories together is that attack economics have inverted. The traditional protection for a small or mid-sized business was never a firewall, it was effort, because a skilled attacker had better things to do than spend two weeks on a company with forty employees. An agent has no better things to do, does not get bored, and costs a few dollars an hour to run. The Strix case shows the second half of the shift, which is that patience is now free, so a credential left in a build artefact in 2023 is not a historical curiosity, it is a live finding waiting for something with infinite time to go looking. For businesses this collapses a comfortable assumption that has quietly underpinned a lot of security budgets. The AEPD case also creates a genuinely novel legal problem, because breach law assumes a controller who failed and an attacker who acted, and it is not yet settled how liability distributes when the acting party is a model, the instructing party wrote one prompt, and the model vendor built guardrails that did not hold.
Implications for People/Careers
The honest advice here is unglamorous, and it is the same advice that has always worked, which is that the things that stop an agent are the things that stop a bored intern: multi-factor authentication everywhere, no shared logins, credentials rotated when people leave, and no secrets living in code, build files, config backups or old repositories. What changes is the deadline, because the gap between a sloppy practice and someone discovering it used to be years and is now closer to hours. Anyone who handles client data should spend twenty minutes this week finding out where their organisation's credentials actually live, and if the answer is partly in a spreadsheet, that is your finding. For those building a career rather than defending one, note that the AEPD's three-origin uncertainty points at a skill nobody currently has enough of, which is being able to reconstruct what an autonomous system did and why. Agent forensics, incident reconstruction and the plain ability to read an audit trail are about to be worth a great deal, and they sit closer to careful investigative thinking than to deep technical training.
Our Future//Take
Our prediction is that within twelve months, agent activity logging becomes a standard clause in commercial contracts and an insurance question, in the same way that breach notification timelines did after GDPR landed. Expect the first serious legal fight over who is liable when an autonomous system causes a breach, and expect it to be messy, because the current frameworks were written for tools that do what they are told rather than systems that decide how to get there. The practical instruction for this week is smaller and it costs nothing. Go and find every place your business still stores a credential, and assume something tireless is already looking for it. The Strix finding was a token from 2023 that nobody remembered writing. Most businesses have one of those. The difference now is that there is something patient enough to find it. Hereβs your βΉ25,000 AI Gift for FREE π

Quick summaries of this week's top AI news, their relevance to your career, and our expert opinions.
Google shipped Gemini 3.8 Live on September 15, aimed at cost-efficient conversational agents, alongside Gemini 3.8 Live Extended Thinking, which reasons and speaks at the same time rather than pausing to think. Extended Thinking took first place on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6, and hit 97.7% on Big Bench Audio. The number worth staring at is a different one: on Sierra's banking voice benchmark, the hardest realistic customer-service test in the set, it scored 35.1%. Both models ship through the Gemini API, AI Studio, Search Live, Gemini Enterprise in private preview, and across Workspace, Gmail and Keep.
Why It Matters to You
Voice is now the most competitive category in AI, and that is good news if you buy rather than build, because the price of a decent voice agent is falling fast. But read those two benchmark numbers together before you promise anything to a customer. Near-perfect on audio comprehension, roughly one in three on a realistic multi-step banking conversation. The model can hear you perfectly and still fail the task, because the difficulty was never the listening, it was holding a goal across a messy conversation where the caller changes their mind.
Our Take
Design for the 35%, not the 82.6. Any voice deployment that matters should be scoped to conversations with a narrow, verifiable outcome, so booking, qualifying, confirming or triaging, and every one of those should have a clean handoff to a human that triggers early rather than after three failed attempts. The teams getting real results from voice agents right now are not the ones with the best model, they are the ones who picked narrow enough jobs. Extended Thinking reasoning while it speaks is the genuinely new capability here, since the long silence while a model thinks is what makes most voice agents feel robotic.
On September 14 Anthropic ended a temporary 50% weekly boost on Claude Code that had run since May, replacing it with what it calls a permanent 25% increase over pre-May limits, which nets out as a 17% cut against the summer allowance for Pro, Max, Team and seat-based Enterprise plans. Free plans and consumption-based Enterprise seats are unaffected. Separately, the HarnessTax study published September 17 benchmarked 21 model-and-tool combinations across seven models and three coding harnesses, and found that harness choice barely moves the success rate but substantially changes the token bill. The same model lands at similar accuracy at very different cost. In fairness, this partly revises a finding we covered a fortnight ago suggesting the expensive option also performed better, and the authors note that genuinely hard problems still benefit from a harness that provides structure.
Why It Matters to You
Two lessons, one uncomfortable. The first is that the capacity you bought is a vendor decision, not a purchase, and a plan's limits can be revised downward between one month and the next. If a business process now depends on an AI subscription, that process has a dependency nobody has priced. The second is more actionable: when the tool wrapped around a model changes your bill without changing your results, you are paying for packaging.
Our Take
Build for substitution. Anything running on a single AI vendor with no tested fallback is one pricing email away from a bad week, and knowing which alternative you would switch to, before you need it, takes an afternoon. On cost, keep asking vendors the question we suggested a fortnight ago, which is what a completed unit of work costs rather than which model is inside. This week's study is a reminder that the answer moves, and that it is worth re-checking quarterly rather than deciding once.
Judge Leonie Brinkema unsealed a 106-page remedies order on September 16 in the Google ad tech antitrust case, rejecting the breakup the Department of Justice had asked for and imposing behavioural remedies instead, effective in 60 days. Google must build API integrations connecting AdX and DFP to open-source Prebid header bidding, must submit AdX bids to rival ad servers on the same terms it gives its own, and must accept a court-appointed technical monitor with full access to its systems and source code for six years, against the fifteen the DOJ requested.
Why It Matters to You
If you buy digital advertising, the plumbing underneath your campaigns is about to change while your dashboards look identical. Forcing Google's exchange to bid into rival ad servers on equal terms is the kind of change that shows up as shifting effective CPMs and altered inventory access rather than as an announcement. Over the next two quarters, comparative performance between channels may move for reasons that have nothing to do with your creative or your targeting.
Our Take
If you buy digital advertising, the plumbing underneath your campaigns is about to change while your dashboards look identical. Forcing Google's exchange to bid into rival ad servers on equal terms is the kind of change that shows up as shifting effective CPMs and altered inventory access rather than as an announcement. Over the next two quarters, comparative performance between channels may move for reasons that have nothing to do with your creative or your targeting. AI Mastery for FREE (Sign up Now).

Discover a comprehensive guide to an AI tool, exploring its features, practical use cases, and learning resources to help you master it.

π Claude
Anthropic merged Claude chat and Cowork into a single interface on September 16, which is the update that makes Claude worth a proper look for non-technical professionals rather than just developers.
The distinction that used to confuse everyone has gone. Chat was the thing in your browser that answered questions and made you copy the result out. Cowork was a separate desktop app that could reach into the folders on your actual computer, work with your real files, and hand you finished documents back. Now it is one product that routes automatically between conversation, artifacts and design work without you choosing a mode. The merge also adds presentation creation with PowerPoint and PDF export, plus collaborative document editing. It is rolling out to Pro and Max subscribers first on web, desktop and mobile, with free and team tiers following.
β Top Features
Works on your real files. The desktop app reads from a folder you nominate, so there is no uploading, no per-file size ceiling, and no twenty-file limit. Point it at a folder of client contracts and ask a question across all of them.
Delivers finished files. Output lands in your folder as a working document, spreadsheet or deck rather than as text you have to copy and reformat.
Presentations with real export. New this week, and the export to PowerPoint is the part that matters, since a deck you cannot open in the software your client uses is not a deck.
Connectors. Gmail, Google Drive, Calendar and others, so it can work across your mail, files and schedule together rather than one at a time.
Automatic routing. It decides whether a request needs a conversation, a document or a design, which removes the main thing that stopped ordinary users getting value out of it.
Skills. Save a repeated instruction set once and reuse it across your work, which is how a good prompt stops being something you rewrite every Monday.
Resources for Learning
Official Documentation: support.claude.com for connectors, file access, plan limits and how Skills work.
The Merge Explained: techcrunch.com for what changed this week and what is rolling out to which tier.

A curated list of noteworthy AI tools and their key details to help you stay ahead in your field.

Launched in New York on September 11, AdAI uses agents to produce finished image and video advertisements from your existing brand assets and guidelines, rather than generating isolated copy or standalone images you then have to assemble. Incubated by Revolution Venture Studios. The differentiator is the complete creative unit, because most AI ad tools hand you pieces and leave the production step to you, which is where the time actually goes. Worth testing if your ad output is limited by creative production rather than by budget. Brand new, so review closely before anything reaches a customer.

Launched September 10 as a native voice agent inside Intermedia's Unite and Contact Center products, built to answer inbound calls, capture opportunities that would otherwise be missed, and route conversations without a human picking up first. The differentiator is that it lives inside the phone system itself rather than bolting on as a separate service, which removes the integration work that usually stalls voice projects at smaller companies. The relevant question for any business evaluating this category remains how much of the follow-up it completes, not how well it answers.

Released September 10, Frames builds working business interfaces from a plain-language prompt: data-entry forms, scenario planners and executive presentations, connected live to Pigment's planning platform rather than sitting as static mockups. The differentiator is that the generated interface is connected to real data, so a scenario planner you describe in a sentence actually calculates. Most useful if you already run planning or forecasting in a dedicated tool and are tired of the gap between the model and the people who need to use it.

A Y Combinator alum launched Atli on September 10, an AI travel companion that lives inside WhatsApp and finds, buys and activates eSIM data across more than 200 countries through a chat conversation. Built by XForge Mobile. The differentiator is that it completes a real transaction end to end, from discovery through payment to activation, inside a messaging app rather than sending you to a website. Beyond the travel use, it is the cleanest example yet of what agentic commerce looks like in a channel Indian customers already live in, which makes it worth studying even if you never need an eSIM.

A quick poll to help you recollect and engage with key points from the newsletter.
Spain's data protection agency filed what it calls the first data breach executed end to end by an AI agent. How much human direction did the attack require after the initial setup? |

Share your feedback on today's edition to help us improve and better meet your needs.
How was Todayβs Edition? |
Share our Newsletter β©
Enjoying our insights on the latest AI breakthroughs? Donβt keep it to yourself! Share this newsletter with friends and colleagues who are passionate about technology and AI innovation.
If you havenβt subscribed yet, make sure to subscribe here to stay updated with cutting-edge AI news, tools, and tutorials delivered straight to your inbox!
Ask Us Anything AI β
Got questions? We've got answers!
Submit your questions, queries, thoughts, opinions or anything regarding AI and we'll feature and address them in our next newsletter. Your curiosity drives our content!
π Reply to this email with your questions, and we'll answer them in our next edition!π

