Every growing Salesforce org runs into the same conversation eventually: storage costs are climbing, and it’s all starting to feel a little sluggish. It’s at this point that someone on the team floats the idea of archiving old records to free up space. So data gets pushed into storage, and everyone moves on to their next initiative.
What almost nobody stops to ask is what actually happened to that data once it left.
Most teams think of that archived data, and honestly their backups in general, as something that just sits there in case of an emergency. But that history is actually one of the most valuable things in the org – most teams just never built it to be truly useful. That’s really the choice teams are up against: whether to archive history out to keep the org fast, or to keep it all and watch the costs climb.
Either way, teams will still end up with data they can’t actually use. Usable and owned storage is the fix, and that’s also what matters when it’s time to do something like run an AI model or pull a report.
Why Backups Get Treated Like an Afterthought
Backups have always been considered part of the “insurance policy” category – necessary, but not a part of the core processes that make an org run better. Luckily, AI has been starting to complicate that assumption, acting faster than most teams expected.
Architects have moved fast here. According to SF Ben’s 2026 Salesforce Architect Survey, AI adoption in the ecosystem is now sitting at 96.7%, up from 88.9% just a year ago, with nearly two-thirds of architects using it daily. That is no longer a pilot, but a day job instead.
That being said, the same survey has found that the barriers holding teams back have changed along with it. Just a year ago, cost was the obvious top concern, but now it is barely in the top five, with the current list of most relevant concerns being:
- Trust in AI outputs (19.3%) – now the single biggest barrier.
- Skills and knowledge gaps (19.0%).
- Accuracy concerns (17.6%).
- Company buy-in (17.3%).
- Cost, now down to just 16.1%.
This data matches what analysts are seeing more broadly, with Gartner predicting that through 2026, organizations are going to abandon 60% of AI projects that are not backed by AI-ready data. Trust and readiness are becoming the real gating factors instead of cost.
In a way, this is a question of governance and reliability: how confident are architects in what the AI is actually working with? A fair amount of that confidence tracks back to the archived, backed-up, “we’ll deal with it later” data that nobody prepared to be used in queries to begin with.
The Cost of Data You Can’t Use
Here’s where the gap actually shows up. In GRAX’s survey of 54 senior enterprise IT and data leaders, 54% said they already have AI running on their SaaS data in some form, but 76% of those same teams also stated that their data is, at best, only partially ready for it.
The mismatch of that size is not small by any account, showing up in orgs that have already invested in at least some form of retention. A common instinct is to reach for the option that is fastest to set up, and all Salesforce teams have several paths available (each with its own trade-off):
- Big Objects keep data native to Salesforce with a relatively low cost in large volumes. There is a limit of 100 objects per org, with indexing rules that have to be decided upfront without the ability to change them later. These objects are not built for ad-hoc reporting and are much better in compliance retention instead of analytics.
- External Objects via Salesforce Connect provides virtualization capabilities for data sitting outside the org, avoiding duplication. Yet, every read operation is a separate live API call, making this method’s performance directly dependent on the external system’s uptime; it is less than ideal for heavy AI or analytics workloads.
- Off-platform export to cloud storage or a data lake offers great scaling and tends to be the cheapest option available when it comes to cost per gigabyte. It also puts the entire burden of structure, indexing, and version history on whatever is managing the target storage.
- Data Cloud aims to unify and ground data for AI and analytics. At its core, this is a harmonization layer that is only as good as the historical data prepared to feed into it. Data Cloud cannot operate as a retention strategy in itself.
None of these choices are inherently wrong on their own. The issue arises when one of these is chosen for either cost or convenience reasoning without weighing what it means for governance, accessibility, or AI-readiness down the line.
This is exactly what data leaders report runs into, with the following reasons being the biggest barriers for getting data AI-ready (according to the same GRAX survey):
- Governance and compliance gaps (69%): data that’s technically stored but not properly tracked or auditable.
- Data quality issues (57%): inconsistent, incomplete, or unreconciled records once they’ve been moved out of the live org.
- Data silos (43%): history split across export formats, tools, and cold storage nobody remembers how to query.
- Access and ownership unclear (33%): nobody’s quite sure who’s responsible for it anymore.
What It Looks Like When Your History Works for You
There is no one right answer here. It is essentially an architectural decision, and the most fitting option is derived from what the data needs to do.
A compliance-driven audit trail that rarely gets queried could feel perfectly fine with a Big Object or a well-indexed off-platform archive. Data that occasionally needs a live lookup from a legacy environment would work well with External Objects.
But data that’s meant to feed AI, forecasting, or cross-system analytics tends to require environments that are built specifically with their purpose in mind: structured, versioned, and centralized enough that a model or a person can actually query the full record history.
That last point is also much more important now than it used to be, with GRAX’s survey finding that 93% of data leaders claim historical data as necessary for AI initiatives. More than a third of those respondents are also saying that the full change history is needed to achieve real value.
Whichever combination of native and external tools a team lands on – that is the bar worth designing toward.
Final Thoughts
The teams that get this right end up with something most Salesforce orgs never have: a full, usable record of their own history that’s ready the moment it’s needed. That’s all to say that archiving to save on storage isn’t the wrong move.
On the contrary, it’s often a necessary one most teams will need to make. But it only solves half the problem if what they’re keeping can’t be queried or fed into an AI model down the line. That gap can be harder to ignore as more teams try to put their historical Salesforce data to work, whether that’s for forecasting or training their AI tools.
The teams getting ahead of this are the ones who can actually put their hands on that old data fast, whether it’s a person asking the question or a model.
If you’re not sure whether your archived Salesforce data could actually support an AI initiative today, that’s worth finding out before you need it. Watch a GRAX demo to see what it takes to make your data usable, not just stored.






