Skip to main content
The catalog importer is our open-source command-line tool for syncing data into your Catalog from anywhere: files in your repositories, Backstage, an internal API, or your data warehouse. Run it on a schedule or from CI, and your catalog types and entries stay in step with the source. Use it when your data doesn’t live in a tool we have a native integration for, or when you want control over exactly which types, entries, and attributes you import. If you’d rather manage a type in a repository without writing the config yourself, the Manage in GitHub button sets up the importer for you. See Managing catalog types in GitHub.

How it works

You describe what to import in a config file, written in Jsonnet, YAML, or JSON. The config holds a sync_id and a list of pipelines. Each pipeline has:
  • Sources: where the data comes from, like a file, a command, or the Backstage API.
  • Outputs: the catalog types to create, and how each source record maps to an entry’s name, external ID, and attributes.
Each time you run sync, the importer makes the catalog types in your config match the source:
  • It creates and updates entries to match what the sources return, identifying each one by its external_id, so renaming an entry doesn’t create a new one.
  • It deletes entries that are no longer in the source. Deleted entries are archived, and one comes back if its external_id reappears in the source. If a source suddenly returns nothing, the importer stops rather than empty the type, unless you pass --allow-delete-all.
  • It only touches the types it owns. Types are tagged with your sync_id, so separate importers, and types you manage in the dashboard, are left alone. Manage each type in one place: if two pipelines produce entries with the same external_id, they overwrite each other on every run, and a type managed with Terraform shouldn’t also be synced by the importer.
  • It leaves types you’ve removed from the config, unless you pass --prune.
A source that fails partway through can return fewer entries than it should, and the importer deletes the rest. Make sure your exec commands exit with an error when they fail, rather than printing partial output, so the sync stops instead.

Taking over a type you built in the dashboard

To start syncing a catalog type you created in the dashboard, give its output the same type_name, and declare every attribute it already has. If the config leaves an existing attribute out, the sync fails. For attributes whose values you still want to edit in the dashboard, like a Team type’s escalation paths, set schema_only: true. See Schema-only attributes.

Get started

  1. Install the importer:
    There are also binaries on the releases page, and a Docker image. Use the latest release: recent versions are much faster on large catalogs.
  2. Create an API key with the View catalog and Manage catalog permissions, and export it:
  3. Generate a starting config, choosing a template for Backstage, Jira Service Management Assets, or files:
  4. Check the config, and preview the changes without making them:
    The dry run prints a diff of every change. Entries it would delete are listed under Deleting unmanaged entries, with each value going to nil or empty.
  5. Run the sync:

Sources

See Sources for every option.

Import from anything with exec

The exec source runs a command and imports whatever JSON or YAML it prints. Use it to load data from somewhere the other sources don’t reach, like an internal API, a query against your data warehouse, or a file that needs reshaping first. This pipeline reshapes a JSON file with jq, and imports each service with its owning team:
The same approach loads a customer list from BigQuery with bq query --format=json, or entries from an internal API with curl. See the exec source for both.
The command runs wherever the importer runs, so the tool it calls needs to be installed there. The Docker image includes only the importer itself, so build your own image on top of it if you need jq, bq, or anything else.

Run it on a schedule

Run the importer from CI, so the Catalog updates whenever the source changes, or on a schedule. The importer’s docs have ready-made configs for GitHub Actions, CircleCI, and GitLab CI. Before you use one, update the importer version it installs to the latest release. Run sync with --dry-run on branches, and only run a real sync, or --prune, from your main branch. Pass --source-repo-url with the URL of the repository that holds your config. The Catalog links each type to that repository and stops it being edited in the dashboard, so changes go through the repository instead.

Learn more

The catalog importer documentation covers: