Docs
Automated Scan Ingestion

Automated Scan Ingestion

Guide for setting up automated scan ingestion from S3 Buckets.

Automated Scan Ingestion Guide

ACE has the capability to automatically ingest scan files from remote blob storage. Without the need to manually upload scans one by one, this can significantly speed up workflow.

This works by making a “Bucket Ingest Configuration” (BIC) and an “Environment Ingest Setting.” (EIS) A BIC is a specific way of connecting to an S3 bucket, using specific credentials, either via Access Key or Assume Role. An EIS is, as implied, per-environment, and tells ACE that a specific environment will use a specific BIC for ingestion. Once an EIS is created, a folder is created in the S3 bucket marked by the environment UUID, and the user can create a folder within for files under a specific tool name. (s3://<bucket-name>/<env-uuid>/<tool-name>/<file-here>)

ACE will periodically pass through the tool folders for scans. If there are files not in a tool name folder (for instance, the root beneath the environment UUID), they are moved to a _skipped folder. While processing each scan report, they will be moved to a dedicated _processing folder underneath the environment UUID path. After completion, they will be moved to a _processed folder, _failed if there was an error during processing, or, if the delete_after_ingest setting is set to true in the BIC, deleted.

Setup

As mentioned before, to use automatic ingestion, one must first create a “Bucket Ingest Configuration” (BIC). A successful connection test (via Assume Role or Access Keys) must be performed before creation. Then tell an environment to ingest using this BIC via an “Environment Ingest Setting” (EIS).

AWS Side

For ACE to ingest from an S3 bucket, a policy must be created for an IAM user or role per bucket. The minimum is listed below:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:ListBucket"],
    "Resource": "arn:aws:s3:::<bucket-name>",
  }, {
    "Effect": "Allow",
    "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
    "Resource": "arn:aws:s3:::<bucket-name>/*"
  }]
}

ListBucket is needed on the root prefix to list contents for ingestion. GetObject is required for granular retrieval of items within the bucket below the root prefix. Because the ingestion pipeline moves files while processing (to folders such as _processing, _processed, _skipped, _failed, etcetera.), PutObject and DeleteObject are required to simulate moving. Furthermore, should the user mark delete_after_ingest true in their Bucket Config creation, the pipeline will also delete scans instead of placing them in _processed.

ACE Side

To create a BIC, navigate to the /admin/settings page. At the bottom should be “Bucket Ingestion Configurations.” One must provide a bucket URI with optional trailing prefixes (this folder must exist to work; ACE only creates a folder within this when an EIS is created), deletion after ingestion, and choose between either the Access Key or Assume Role authorization methods for contacting the bucket. Then, they must perform a valid Connection Test to be able to create the configuration. This allows for explicit separation between different configurations across environments.

To create an EIS, navigate to the Environment Page, then to the Sources tab. One can then assign a BIC to this environment, manually search for scans within the bucket, check past logs and reports on ingestion, and disable scanning.

If during an automatic run, the connection fails, the BIC is disabled. To reenable it, one must updated the credentials and run a successful connection test.