Your SaaS Vendors Are Training AI on Your Data. Here Is How to Check
September 16, 2026
Over the past year, a pattern has emerged across major SaaS providers. AI features get announced, and somewhere in the accompanying policy update, the vendor reserves the right to use customer data to train its models. The setting that governs this is typically enabled by default, and the burden of turning it off falls on the customer.
This is worth your attention for a simple reason. The data these platforms hold, including your code, your customer records, your internal documentation, and your team's conversations, is among the most valuable material your company owns. Whether a vendor gets to learn from it should be a decision someone in your organization makes deliberately, not a default nobody noticed.

This article documents what three major vendors have announced, what the policies actually say, and how to check your own settings.
GitHub: Copilot Interaction Data, Opt-Out by Default
In March 2026, GitHub announced an update to its Copilot data usage policy. Starting April 24, 2026, interaction data from Copilot Free, Pro, and Pro+ users is used to train and improve GitHub's AI models unless the user opts out.
Two details matter here. The first is the scope: the policy covers individual plans, while Copilot Business and Enterprise remain excluded from training by default. The second is the mechanism: existing users were transitioned to the new policy on a set date, with opt-out available in the Copilot settings of the user account. Users who did not act by that date started contributing interaction data automatically.
The full announcement is on the GitHub blog, and the current policy details are in GitHub's documentation.
Atlassian: The Data Contribution Program, Effective August 17
Atlassian introduced what it calls the "data contribution" program, documented on its trust pages. The changes took effect on August 17, 2026, following a rollout of new data contribution settings that began in April and completed in May.
The framing deserves a careful read. "Contribution" describes a transfer of value from the customer to the vendor: your Jira issues, Confluence pages, and support tickets become material Atlassian uses to improve its apps for all customers. Whether that trade is acceptable depends on your data and your industry, but it is a trade, and it should be evaluated as one.
The default settings vary by plan tier and are worth knowing exactly. Metadata contribution is on by default for all plans, including Free, Standard, Premium, and Enterprise. In-app data contribution is on by default for Free and Standard, and off by default for Premium and Enterprise. Only Enterprise customers can opt out of metadata contribution, while everyone can adjust the in-app data setting. Atlassian describes safeguards including de-identification and aggregation before contributed data is used, and initially the settings apply to Jira, Confluence, and Jira Service Management, along with data in Atlassian Platform apps.
For Atlassian administrators, the participation settings are managed at the organization level in Atlassian Administration. The trust page above documents what is collected, how it is protected, and where the controls live.
HubSpot: A Reversal Worth Studying
HubSpot's case is instructive because of how it ended. On July 1, 2026, HubSpot announced a terms of service update tied to a new prospecting feature called Contact Discovery, which involved a shared enrichment dataset of professional contact details. The reaction from the community was significant, and four days later HubSpot published an official response titled "We Got This Wrong. And We Are Fixing It", confirming that the terms of service changes would not move forward and reaffirming that "your CRM data belongs to you."
The episode demonstrates two things. Vendors do respond when customers push back collectively and publicly, which is a reason to pay attention and speak up rather than assume these decisions are final. And the specific commitment HubSpot made in its response is worth quoting for future reference: any new enrichment capabilities that make use of customer data will be "fully and transparently opt-in." That is the standard the industry should be moving toward, and one to hold vendors accountable for going forward.

Why Opt-Out Defaults Matter
Each vendor's policy, read individually, includes reasonable-sounding protections: aggregation, de-identification, filtering of sensitive content. The structural issue is not any single safeguard but the direction of the default.
An opt-in default means the vendor must convince you the trade is worth it. An opt-out default means the vendor collects unless someone in your organization knows the setting exists, has the permissions to change it, and acts before the deadline. Across thousands of customers, the difference between those two defaults is enormous, and vendors know it. That is precisely why opt-out is becoming the industry pattern.
There is also a compounding effect with a topic we have covered before. The more a vendor's AI features are trained on the data inside its ecosystem, the more the product's value depends on your data staying there, which deepens the dynamics described in our guide to vendor lock-in.
How to Audit Your Own Stack

A practical review takes an afternoon and is worth repeating twice a year. For each SaaS platform your company relies on:
- Search the vendor's trust center or privacy documentation for terms like "model training," "AI improvement," "data contribution," or "machine learning."
- Identify whether the training setting is opt-in or opt-out, and what its current state is for your organization.
- Check which plan tiers the policy applies to. Enterprise tiers are often excluded by default while individual and lower tiers are included.
- Verify who in your organization has the admin permissions to change the setting, and document the decision either way.
- Set a reminder to re-check after major product announcements, because these policies change with little notice.
The goal is not necessarily to opt out of everything. Some trades may be acceptable to your organization. The goal is that the decision gets made by someone accountable, with the policy actually read, rather than by a default.
FAQ
Do SaaS vendors use customer data to train AI models?
An increasing number do, typically under policies enabled by default with an opt-out mechanism. Documented examples include GitHub's Copilot interaction data policy for individual plans and Atlassian's data contribution program. The specifics vary by vendor, plan tier, and data type, so the reliable answer for your stack requires checking each vendor's current documentation.
How do I stop a vendor from training AI on my company data?
Look for the relevant setting in the platform's admin or account configuration, usually documented in the vendor's trust center under terms like "data contribution," "model training," or "AI data usage." Note that permissions matter: in most platforms, only organization administrators can change these settings for the whole company. Note also that not every setting is available on every plan: Atlassian, for example, only allows Enterprise customers to opt out of metadata contribution.
Is my data protected even if I do not opt out?
Vendors typically describe safeguards such as de-identification, aggregation, and filtering of secrets or sensitive content. How effective these are varies and is difficult to verify from outside. The practical position is to treat the safeguards as risk reduction rather than elimination, and to make the opt-in or opt-out decision based on the sensitivity of the data the platform holds.
Why are these settings enabled by default?
Because defaults determine outcomes at scale. Most customers never review these settings, so an opt-out default yields far more training data than an opt-in default would. Regulatory pressure and customer pushback, as in HubSpot's July 2026 reversal, are currently the main forces pushing vendors toward more explicit consent.