Profile Cloud SQL data in a single project

This page describes how to configure Cloud SQL data discovery at the project level. If you want to profile an organization or folder, see Profile Cloud SQL data in an organization or folder.

For more information about the discovery service, see Data profiles.

How it works

The following is a high-level workflow for profiling Cloud SQL data:

  1. Create a scan configuration.

    After you create a scan configuration, Sensitive Data Protection starts identifying your Cloud SQL instances and creating a default connection for each instance. Depending on the number of instances in scope of discovery, this process can take a few hours. You can exit the Google Cloud console and check your connections later.

  2. Grant the required IAM roles to the service agent associated with your scan configuration.

  3. When the default connections are ready, give Sensitive Data Protection access to your Cloud SQL instances by updating each connection with the proper database user credentials. You can provide existing database user accounts or create database users.

  4. Recommended: Increase the maximum number of connections that Sensitive Data Protection can use to profile your data. Increasing the connections can speed up discovery.

Supported services

This feature supports the following:

  • Cloud SQL for MySQL
  • Cloud SQL for PostgreSQL

Cloud SQL for SQL Server isn't supported.

Processing and storage regions

Sensitive Data Protection is a regional and multi-regional service; it doesn't distinguish between zones. When Sensitive Data Protection profiles a Cloud SQL instance, the data is processed in its current region, but not necessarily its current zone. For example, if a Cloud SQL instance is stored in the us-central1-a zone, then Sensitive Data Protection processes and stores the data profiles in the us-central1 region.

For more information, see Data residency considerations.

Before you begin

  1. If you have an organization-level discovery subscription—including one through Security Command Center—be aware that this project-level discovery configuration isn't included in your subscription and is billed separately. We recommend that you use an organization-level discovery configuration to profile the project. For more information, see Profile select projects or data assets in an organization or folder.

  2. Make sure the Cloud Data Loss Prevention API is enabled on your project:

    1. Sign in to your Google Cloud account. If you're new to Google Cloud, create an account to evaluate how our products perform in real-world scenarios. New customers also get $300 in free credits to run, test, and deploy workloads.
    2. In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

      Roles required to select or create a project

      • Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
      • Create a project: To create a project, you need the Project Creator role (roles/resourcemanager.projectCreator), which contains the resourcemanager.projects.create permission. Learn how to grant roles.

      Go to project selector

    3. Verify that billing is enabled for your Google Cloud project.

    4. Enable the required API.

      Roles required to enable APIs

      To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

      Enable the API

    5. In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

      Roles required to select or create a project

      • Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
      • Create a project: To create a project, you need the Project Creator role (roles/resourcemanager.projectCreator), which contains the resourcemanager.projects.create permission. Learn how to grant roles.

      Go to project selector

    6. Verify that billing is enabled for your Google Cloud project.

    7. Enable the required API.

      Roles required to enable APIs

      To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

      Enable the API

  3. Confirm that you have the IAM permissions that are required to configure data profiles at the project level.

  4. You must have an inspection template in each region where you have data to be profiled. If you want to use a single template for multiple regions, you can use a template that is stored in the global region. If organizational policies prevent you from creating an inspection template in the global region, then you must set a dedicated inspection template for each region. For more information, see Data residency considerations.

    This task lets you create an inspection template in the global region only. If you need dedicated inspection templates for one or more regions, you must create those templates before performing this task.

  5. You can configure Sensitive Data Protection to send notifications to Pub/Sub when certain events occur, such as when Sensitive Data Protection profiles a new table. If you want to use this feature, you must first create a Pub/Sub topic.

  6. You can configure Sensitive Data Protection to automatically attach tags to your resources. This feature lets you conditionally grant access to those resources based on their calculated sensitivity levels. If you want to use this feature, you must first complete the tasks in Control IAM access to resources based on data sensitivity.

  7. You can configure Sensitive Data Protection to automatically attach aspects to the profiled Cloud SQL resources based on discovery insights. If you want to use this feature, you must enable the integration of universal catalog on your Cloud SQL for MySQL instance or Cloud SQL for PostgreSQL instance.

Create a scan configuration

  1. Go to the Create scan configuration page.

    Go to Create scan configuration

  2. Go to your project. On the toolbar, click the project selector and select your project.

The following sections provide more information about the steps in the Create scan configuration page. At the end of each section, click Continue.

Select a discovery type

Select Cloud SQL.

Select scope

Do one of the following:

  • If you want to scan a single table, select Scan one table.

    For each table, you can have only one single-resource scan configuration. For more information, see Profile a single data resource.

    Fill in the details of the table that you want to profile.

  • If you want to perform standard project-level profiling, select Scan selected project.

Manage schedules

If the default profiling frequency suits your needs, you can skip this section of the Create scan configuration page.

Configure this section for the following reasons:

  • To make fine-grained adjustments to the profiling frequency of all your data or certain subsets of your data.
  • To specify the tables that you don't want to profile.
  • To specify the tables that you don't want profiled more than once.

To make fine-grained adjustments to profiling frequency, follow these steps:

  1. Click Add schedule.
  2. In the Filters section, you define one or more filters that specify which tables are in the schedule's scope. A table is considered to be in the schedule's scope if it matches at least one of the filters defined.

    To configure a filter, specify at least one of the following:

    • A project ID or a regular expression that specifies one or more projects.
    • An instance ID or a regular expression that specifies one or more instances.
    • A database ID or a regular expression that specifies one or more databases.
    • A table ID or a regular expression that specifies one or more tables. Enter this value in the Database resource name or regular expression field.

    Regular expressions must follow RE2 syntax.

    For example, if you want all tables in a database to be included in the filter, enter the database ID in the Database ID field.

    To match a filter, a table must meet all the regular expressions specified within that filter.

    If you want to add more filters, click Add filter and repeat this step.

  3. Click Frequency.

  4. In the Frequency section, specify whether the discovery service should profile the tables you selected and, if so, how often:

    • If you never want the tables to be profiled, turn off Do profile this data.

    • If you want the tables to be profiled at least once, leave Do profile this data on.

      In the succeeding fields in this section, you specify whether the system should reprofile your data and what events should trigger a reprofile operation. For more information, see Frequency of data profile generation.

      1. For On a schedule, specify how often you want the the tables to be reprofiled. The tables are reprofiled regardless of whether they underwent any changes.
      2. For When schema changes, specify how often Sensitive Data Protection should check if the selected tables had schema changes after they were last profiled. Only tables with schema changes will be reprofiled.
      3. For Types of schema change, specify which types of schema changes should trigger a reprofile operation. Select one of the following:
        • New columns: Reprofile the tables that gained new columns.
        • Removed columns: Reprofile the tables that had columns removed.

        For example, suppose you have tables that gain new columns every day, and you need to profile their contents each time. You can set When schema changes to Reprofile daily, and set Types of schema change to New columns.

      4. For When inspect template changes, specify whether you want your data to be reprofiled when the associated inspection template is updated, and if so, how often.

        An inspection template change is detected when either of the following occurs:

        • The name of an inspection template changes in your scan configuration.
        • The updateTime of an inspection template changes.

      5. For example, if you set an inspection template for the us-west1 region and you update that inspection template, then only data in the us-west1 region will be reprofiled.

  5. Click Conditions.

    In the Conditions section, you specify the types of database resources that you want to profile. By default, Sensitive Data Protection is set to profile all supported database resource types. When Sensitive Data Protection adds support for more database resource types, those types will automatically be profiled, too.

  6. Optional: If you want to explicitly set the database resource types that you want to profile, follow these steps:

    1. Click the Database resource types field.
    2. Select the database resource types that you want to profile.

    If Sensitive Data Protection later adds discovery support for more Cloud SQL database resource types, those types will only be profiled if you return to this list and select them.

  7. Click Done.

  8. Optional: To add more schedules, click Add schedule and repeat the previous steps.