Thanks to visit codestin.com
Credit goes to docs.datafold.com

Skip to main content
NOTE: Datafold needs catalog-level permissions in your Databricks workspace to read and write table data, stage cross-database diff transfers, query system tables, and deploy migration bundles. You will need workspace admin access to create a service principal and grant the required permissions.
If your workspace storage account is firewalled to block public access, see Securing Connections → Azure. Query results over 1 MB are downloaded directly from workspace storage via Cloud Fetch, which typically requires its own private endpoint in addition to the SQL warehouse connection.
Steps to complete:
  1. Create a service principal and configure authentication
  2. Retrieve SQL warehouse connection details
  3. Grant permissions
  4. Configure your data connection in Datafold
If you plan to run Datafold Migration Agent inference on your own Databricks model serving endpoints, there are two extra prerequisites — see Databricks AI for migrations.

Create a service principal and configure authentication

Create a dedicated service principal for the Datafold integration. This is the identity Datafold will use to connect to your workspace.
  1. Go to SettingsIdentity and accessService principals
  2. Click Add service principal and give it a name (e.g., datafold)
  3. Select the service principal, go to the Secrets tab, and click Generate secret
  4. Save the Client ID and Secret — the secret is only shown once
OAuth secrets are valid for up to 730 days. You can have a maximum of 5 active secrets per service principal. Rotate secrets before expiry to avoid connection interruptions.
Datafold also supports Personal Access Tokens as an alternative authentication method. PATs are considered legacy by Databricks — see the Databricks authentication documentation for details.

Retrieve SQL warehouse connection details

Navigate to SQL Warehouses under the SQL section in the left sidebar. Choose the preferred warehouse and copy the following fields from its Connection Details tab:
  • Server hostname
  • HTTP path
You also need to grant the service principal access to the SQL warehouse:
  1. On the warehouse page, click the Permissions tab
  2. Add the service principal and grant Can Use permission

Grant permissions

Run the following SQL statements to grant Datafold the permissions it needs. Replace <catalog_name> and <service_principal_id> with your values. Replace <schema_name> with the schema where you want to store the DMA bundle volume (e.g., default).
The <service_principal_id> is the application ID (also called Client ID) of your service principal. In Databricks SQL, service principal identifiers must be enclosed in backticks.

(Optional) Additional grant for migrations

If you use Datafold to run migrations (its agents deploy and run Databricks Asset Bundles on this connection), also grant permission to create schemas. Each migration run materializes into a fresh schema created at deploy time. This is not needed for data diffing, monitoring, or lineage.
Authentication for migrations depends on the method. A service principal using M2M OAuth needs no additional token scope. If you use a Personal Access Token (legacy) instead, it must have the all-apis scope, because bundle deployment uploads files via the Databricks workspace-files API, which Databricks gates on all-apis even when narrower scopes are present.

Databricks AI for migrations

This section applies only if the Datafold Migration Agent will run inference on your Databricks model serving endpoints. Skip it for data diffing, monitoring, lineage, and CI, and for migrations that use Datafold-managed inference or a different LLM provider.
Databricks-served inference reuses the data connection above — there is no second service principal, endpoint URL, or API key to create. Two things still need arranging on the Databricks side.

Grant Can Query on the serving endpoints

Serving endpoint access is a separate ACL from SQL warehouse and Unity Catalog access, so the grants above do not cover it. For each endpoint you approve for inference:
  1. Navigate to Serving under the Machine Learning section in the left sidebar
  2. Open the endpoint and click its Permissions tab
  3. Add the Datafold service principal and grant Can Query
Inference requires the connection to authenticate as a service principal — M2M OAuth, a Personal Access Token, or Azure Entra ID. Per-user OAuth connections cannot be used, because the agent runs without a user context.
If your workspace serves pay-per-token foundation models as Unity Catalog model services (system.ai.*), Datafold lists them from the Unity Catalog model services API. When the service principal cannot read that API, those models are missing from the model picker; endpoints already configured keep serving inference.

Raise the AI Gateway throughput limit

The default AI Gateway limit is roughly 200,000 input tokens per minute, which throttles a real migration heavily. Ask your Databricks account team to raise it for the models you approve — request at least 2–3M input tokens per minute. Tier 2 allows up to 10M.
Raising the limit is free: the default is a guardrail, not a paid tier. Databricks typically takes a few days to apply the change, so request it early.
Once Can Query is granted and the limit is raised, connect the endpoints in Datafold by adding the Databricks LLM provider integration.

Configure in Datafold

Select M2M OAuth / Service Principal (Recommended) as the authentication method and fill in the following fields: Click Create. Your data connection is ready!

Validate your setup

Run these queries to verify that permissions are configured correctly: