# Introduction

**Illumina® BioInsight Platform Core** (previously known as Illumina Connected Analytics) is a cloud-based software platform intended to be used to manage, analyze, and interpret large volumes of multi-omics data in a secure, scalable, and flexible environment. The versatility of the system allows the platform to be used for a broad range of applications.

[Get Started](/get-started/gs-getstarted) gives an overview of how to access and configure Platform Core with [Network settings](/reference/r-networksettings) showing the access prerequisites.

### New

To see **what is new** in the latest version, use the [software release notes](/reference/software-release-notes) and the [document revision history](/reference/r-documentrevisionhistory).

### Home

The home section provides an overview of the main Platform Core sections such as [projects](/home/h-projects) (the main work location), [bundles](/home/h-bundles) (asset packages), [logging](/home/h-eventlog), [metadata](/home/h-metadatamodels) (to capture additional information), [Docker](/home/h-dockerrepository) and [tool](/home/h-toolrepository) images (containerised applications) and how to configure (your own) [storage](/home/h-storage).

### Project

**Projects** are your primary work locations which contain your [data](/project/p-data) and [samples](/project/p-samples). Here you will create [**pipelines**](/project/p-flow/f-pipelines) and use them for [**analyses**](/project/p-flow). You configure who can access your project by means of the [teams](/project/p-team) settings. The results can be processed with the help of [Base](/project/p-base), [Bench](/project/p-bench) or [Cohorts](/project/p-cohorts). Projects can be considered as a binder for your work and information.

### Command-Line Interface

This section contains information on how to [download](/command-line-interface/cli-releasehistory), [install](/command-line-interface/cli-installation), [authenticate](/command-line-interface/cli-authentication), [configure](/command-line-interface/cli-configsettings) and [use](/command-line-interface/cli-indexcommands) the command line interface, as alternative for the Platform Core GUI.

### Sequencer Integration

Information and tutorials on [cloud analysis auto launch](/sequencer-integration/analysis_autolaunch).

### Tutorials

There is a set of step-by-step **Tutorials**

* [Nextflow](/tutorials/nextflow)
* [CWL CLI](/tutorials/cli-cwl)
* [Base](/tutorials/base_basics)
* [Bench](/project/p-bench)
* [API](/tutorials/api-introduction)
* [DRAGEN by CLI](broken://pages/1zoggcjo24Y2HsjoINyk)
* [Data Transfer](/tutorials/datatransfer)
* [Pipeline Chaining](/tutorials/pipeline_chaining_aws)
* [DRAGEN Analysis](/tutorials/end-to-end-1)

### Reference

In the **Reference** section, you can find more information on the [**API**](/reference/r-api), [**Pricing**](/reference/r-pricing), [**Security and Compliance**](/reference/r-securityandcompliance),

For an overview of the available subscription tiers and functionality, please refer to [this page](https://emea.illumina.com/products/by-type/informatics-products/connected-analytics.html) on the Illumina website.


# About the Platform

**Illumina® BioInsight Platform Core** (previously known as Illumina Connected Analytics) is a cloud-based software platform intended to be used to manage, analyze, and interpret large volumes of multi-omics data in a secure, scalable, and flexible environment. The versatility of the system allows the platform to be used for a broad range of applications.

When using the applications provided on the platform for diagnostic purposes, it is the responsibility of the user to determine regulatory requirements and to validate for intended use, as appropriate.

The platform is hosted in regions listed below.

| Region Name          | Region Identifier |
| -------------------- | ----------------- |
| Australia            | AU                |
| Canada               | CA                |
| Germany              | EU                |
| India                | IN                |
| Indonesia            | ID                |
| Japan                | JP                |
| Singapore            | SG                |
| South Korea          | KR                |
| United Kingdom       | GB                |
| United Arab Emirates | AE                |
| United States        | US                |

The platform hosts a suite of RESTful HTTP-based application programming interfaces (APIs) to perform operations on data and analysis resources. A web application user-interface is hosted alongside the API to deliver an interactive visualization of the resources and enables additional functionality beyond automated analysis and data transfer. Storage and compute costs are presented via usage information in the account console, and a variety of compute resource options are specifiable for applications to fine tune efficiency.

{% hint style="info" %}
Our systems are synchronized using a Cloud Time Sync Service to ensure accurate timekeeping and consistent log timestamps.
{% endhint %}

## Getting Started

The user documentation provides material for learning the basics of interacting with the platform including examples and tutorials. Start with the [Get Started](/get-started/gs-getstarted) documentation to learn more.

## Getting Help

Use the search bar on the top right to navigate through the help docs and find specific topics of interest.

If you have any questions, contact Illumina Technical Support by phone or email:

Illumina Technical Support | <techsupport@illumina.com> | 1-800-809-4566

For customers outside the United States, Illumina regional Technical Support contact information can be found at [www.illumina.com/company/contact-us.html](http://www.illumina.com/company/contact-us.html).

To see the current Platform Core version you are logged in to, click your username found on the top right of the screen and then select **About**.

## Other Illumina Products

To view a list of the products to which you have access, select the 9 dots symbol at the top right of Platform Core. This will list your products. If you have multiple regional applications for the same product, the region of each is shown between brackets.

The **More Tools** category presents the following options

* My Illumina Dashboard to monitor instruments, streamline purchases and keep track of upcoming activities.
* Link to the Support Center for additional information and help.
* Link to the order management from where you can keep track of your current and past orders.

## Release Notes

In the [Release Notes ](/reference/software-release-notes)section of the documentation, posts are made for new versions of deployments of the core platform components.


# Get Started

## Software Registration

If you are a new user, please consult the [Illumina BioInsight Platform Registration Guide](https://help.connected.illumina.com/account-management/rg-registration) for detailed guidance on setting up an account and registering a subscription.

## Tenant Setup

The platform requires a provisioned tenant in the[ **Illumina account management**](https://help.connected.illumina.com/account-management/admin-console) ([IAM](https://help.connected.illumina.com/account-management/admin-console)) system with access to the **Illumina BioInsight Platform Core** application. Once a tenant has been provisioned, a tenant administrator will be assigned. The **tenant administrator** has permission to manage account access including adding users, creating workgroups, and adding additional tenant administrators.

Each tenant is assigned a domain name used to login to the platform. The domain name is used in the login URL to navigate to the appropriate login page in a web browser. The login URL is `https://<domain_name>.login.illumina.com` with `<domain_name>` replaced by the actual domain name.

For more details on identity and access management, please see the [Illumina BioInsight Platform](https://help.connected.illumina.com/) help.

{% hint style="info" %}
If you have intrusion detection systems active on your infrastructure, be aware that activities performed by Platform Core on your behalf (such as accessing [your own S3 storage](/home/h-storage/s-awss3) ) might trigger suspicious activity alerts. Please review the alerts and rules with your vendor to set up appropriate policies on your detection system.
{% endhint %}

* by the **tenant administrator** by logging in to their domain and navigating to **Illumina Account Management** under their profile at the top right
* or by the **user** by accessing `https://platform.login.illumina.com` and selecting the option **Don't have an account**.

Once the account has been added to the domain, the tenant administrator may assign registered users to [workgroups](https://help.connected.illumina.com/account-management/admin-console/workgroups) which bundle users with permission to use the Platform Core application. Registered users can be made workgroup administrators by tenant administrators or existing workgroup administrators.

## API Keys

If you want to use the [**command-line interface**](https://github.com/illumina-swi/ica-docs/blob/unicorn/docs/get-started/broken-reference/README.md) (CLI) or the [**Application Programming Interface**](/reference/r-api) (API), you can use an [API Key ](https://help.connected.illumina.com/account-management/platform-home)as credentials when logging in. API Keys operate similar to a user name and password combination and must be **kept secure** and **rotated on a regular basis** (preferably yearly). \`

When **keys are compromised or no longer in use, they must be revoked**. This is done through the [domain login URL](https://ilmn.login.illumina.com/platform-home/#/home) by navigating to the User menu item on the left and selecting "API Keys", followed by selecting the key and using the trash icon next to it.

## Generate an API Key

API Keys are limited to 10 per user and are managed through the product dashboard after logging in through the [domain login URL](https://ilmn.login.illumina.com/platform-home/#/home). See[ Managing API Keys](https://help.connected.illumina.com/account-management/platform-home#manage-api-keys) for more information.

{% hint style="warning" %}
For security reasons, do **not use accounts with administrator level access** to generate API keys. Create a specific CLI user with basic permissions instead. This will minimize the possible impact of compromised keys.
{% endhint %}

{% hint style="warning" %}
Once the API key generation window is closed, the key contents will not be accessible through the domain login page, so be sure to store it securely for future reference.
{% endhint %}

***

## Access via Web UI

The web application provides a visual user interface (UI) for navigating resources in the platform, managing projects, and extended features beyond the API. To access the web application, navigate to the [Illumina BioInsight Platform Core portal](https://ica.illumina.com/ica).

* On the **left**, you have the **navigation bar** (1) which will auto-collapse on smaller screens. To collapse it, use the **double arrow symbol** (2). When collapsed, use the >> symbol to expand it.
* The **central** **part** (3) of the display is the item on which you are **performing your actions** and the **breadcrumb** **menu** (4) to return to the projects overview or a previous level. You can also use your browser's back button to return to the level from which you came.
* At the top **right**, you have icons to **refresh contents** (5) I**llumina product access** (6), access to the **online help** (7) and **user** **information** (8).

<figure><img src="/files/EEyX9CVaEI2gEySvNwET" alt=""><figcaption></figcaption></figure>

## Access via the CLI

The command-line interface offers a developer-oriented experience for interacting with the APIs to manage resources and launch analysis workflows. Find instructions for using the command-line interface including download links for your operating system in the [CLI documentation](/command-line-interface/cli-installation).

## Access via the API

The HTTP-based application programming interfaces (APIs) are listed in the [API Reference](/reference/r-api) section of the documentation. The reference documentation provides the ability to call APIs from the browser page and shows detailed information about the API schemas. HTTP client tooling such as Postman or cURL can be used to make direct calls to the API outside of the browser.

{% hint style="info" %}
When accessing the API using the API Reference page or through REST client tools, the `Authorization` header must be provided with the value set to `Bearer <token>` where `<token>` is replaced with a valid JSON Web Token (JWT). For generating a JWT, see [JSON Web Token (JWT)](#json-web-token-jwt).
{% endhint %}

***

## Object Identifiers

The object data models for resources that are created in the platform include a unique `id` field for identifying the resource. These fixed machine-readable IDs are used for accessing and modifying the resource through the API or CLI, even if the resource name changes.

## JSON Web Token (JWT)

Accessing the platform APIs requires authorizing calls using JSON Web Tokens (JWT). A JWT is a standardized trusted claim containing authentication context. This is a primary security mechanism to protect against unauthorized cross-account data access.

A JWT is generated by providing user credentials (API Key or username/password) to the token creation endpoint. Token creation can be performed using the API directly or the CLI.


# Projects

## Introduction

When looking at the main Platform Core navigation, you will see the following structure:

* **Projects** are your primary work locations which contain your data and tools to execute your analyses. Projects can be considered as a binder for your work and information. You can have data contained within a project, or you can choose to make it shareable between projects.
* **Reference Data** are reference genome sets which you use to help look for deviations and to compare your data against.
* **Bundles** are packages of assets such as sample data, pipelines, tools and templates which you can use as a curated data set. Bundles can be provided both by Illumina and other providers, and you can even create your own bundles. You will find the Illumina-provided pipelines in bundles.
* **Audit/Event Logs** are used for audit purposes and issue resolving.
* **System Settings** contain general information susch as the location of storage space, docker images and tool repositories.

Projects are the main dividers in Platform Core. They provide an access-controlled boundary for organizing and sharing resources created in the platform. The Projects view is used to manage projects within the current tenant.

{% hint style="info" %}
There is a combined limit of 30,000 projects and bundles per tenant.
{% endhint %}

## Create new Project

To create a new project, click the **Projects > + Create** button.

On the project creation screen, add information to create a project. See [Project Details](/project/p-details) page for information about each field.

Required fields include:

* **Name**
  * 1-255 characters
  * Must begin with a letter
  * Characters are limited to alphanumerics, hyphens, underscores, and spaces
* **Project Owner** Owner (and usually contact person) of the project. The project owner has the same rights as a project administrator, but can not be removed from a project without first assigning another project owner. This can be done by the current project owner, the tenant administrator or a project administrator of the current project. Reassignment is done at **Projects > your\_project > Project Settings > Team > Edit**.
* **Region** Select your project location. Options available are based on Entitlement(s) associated with purchased subscription.
* **Analysis Priority** (Low/Medium(default)/High) This is balanced per tenant with high priority analyses started first and the system progressing to the next lower priority once all higher priority analyses are running. Balance your priorities so that lower priority projects do not remain waiting for resources indefinitely.
* **Billing Mode** This determine who pays for the project costs (compute and storage) When set to **project**, the tenant of the **project owner will be charged** for compute and storage. When set to **tenant**, the tenant of the **user who started the analysis will be charged** for compute and storage.
* **Data Sharing** Enable this if you want to allow the **data from this project to be linked, moved or copied** and used in other projects of your tenant. Disabling this is a convenient way to prevent your data from showing up in the list of available data to be linked, moved or copied in other projects. *Even though this prevents copying and linking files and folders, it does not protect against someone downloading the files or copying the contents of your files from the viewer.*
* **Storage Bundle** This is auto-selected and appears when you select the Project Region.

<figure><img src="/files/nR9lq8dVW8WdxzQrR1zQ" alt=""><figcaption></figcaption></figure>

Click the **Save** button to finish creating the project. The project will be visible from the Projects view.

{% hint style="info" %}
You may see projects with an additional Information field, this is for backwards compatibility as the field has been superseded by the Short description field.
{% endhint %}

## Create with Storage Configuration

Refer to the [Storage Configuration](/home/h-storage) documentation for details on creating a storage configuration.

During project creation, select the *I want to manage my own storage* checkbox to use a Storage Configuration as the data provider for the project.

With a storage configuration set, a project will have a 2-way sync with the external cloud storage provider: any data added directly to the external storage will be synchronized into the Platform Core project data, and any data added to the project will be synchronized into the external cloud storage.

If there is an issue with your storage configuration, it will be indicated on the project tile.

<figure><img src="/files/Z3zZKQ1ixujCDD77NBoz" alt="" width="372"><figcaption></figcaption></figure>

## Managing Projects

Several tools are available to assist you with keeping an overview of your projects. These filters work in both list and tile view and persist across sessions.

<figure><img src="/files/AuV0TF78oW37UMpAtMeh" alt=""><figcaption></figcaption></figure>

1. **Searching** is a case-insensitive wildcard filter. Any project which contains the characters will be shown. Use \* as wildcard in searches. Be aware that operators without search words are blocked and will result in *Unexpected error occurred when searching for projects*. You can use the brackets, AND, OR and NOT operators, provided that you do not start the search with them (*Monkey AND Banana* is allowed, *AND Aardvark* by itself is invalid syntax)
2. **Filter by Workgroup** : Projects in Platform Core can be accessible for different [workgroups](/project/p-team). This drop-down list allows you to filter projects for specific workgroups. To reset the filter so it displays projects from all your workgroups, use the x on the right which appears when a workgroup is selected.
3. **Hidden projects** : You can hide projects (**Projects > your\_project > Project settings > Details > Hide**) which you no longer use. Hiding will **delete data in base and bench and will thus be irreversible**.
   * You can still see hidden projects if you select this option and delete the data they contain at **Projects > your\_project > Data** to save on storage costs.
   * If you are using **your own S3** bucket, your S3 storage **will be unlinked from the project**, but the data will remain in your S3 storage. **Your S3 storage can then be used for other projects**.
   * Hiding projects is not possible for[ externally-managed](#externally-managed-projects) projects.

<figure><img src="/files/LAvq5Au5vxn6fXsurGsO" alt="" width="364"><figcaption></figcaption></figure>

4. **Favorites** : By clicking on the star next to the project name in the tile view, you set a project as favorite. You can have multiple favorites and use the Favorites checkbox to only show those favorites. This prevents having too many projects visible.
5. **Tile view** shows a grid of projects. This view is best suited if you only have a few projects or have filtered them out by creating favorites. A single click will open the project.
6. **List view** shows a list of projects. This view allows you to add additional filters on name, description, location, user role, tenant, size and analyses. Click on the project name to open it.

{% hint style="info" %}
Items which are shown in list view have an **Export** option at the bottom of the screen. You can choose to support the entire page or only the selected rows in CSV, JSON or Excel format.
{% endhint %}

In **tile view**, your project tiles will show the project name, location, tenant, size, number of analyses and your role.

<div align="center"><figure><img src="/files/81xt79JBDqVS0KWjUoDc" alt="" width="368"><figcaption><p>project tile view</p></figcaption></figure></div>

In **list view**, the star indicates your favourites while warnings, errors and information icons are displayed next to the project name. Hover over those icons to see the details.

<figure><img src="/files/DHVorvLz9OqO3un5Wyio" alt=""><figcaption></figcaption></figure>

<details>

<summary>If you are missing Projects</summary>

If you are missing projects, especially those been created by other users, the workgroup filter might still be active. Clear the filter with the x to the right. You can verify the list of projects to which you have access with the [CLI command](/command-line-interface/cli-indexcommands#icav2-projects) `icav2 projects list`.

</details>

## Externally-managed projects

Illumina software applications which do their own data management on Platform Core (such as BSSH) store their resources and data in a project in the same was as manually created projects work in Platform Core. For Platform Core, these projects are considered **externally-managed projects** and there are a number of **restrictions** on which actions are allowed on externally-managed projects from within Platform Core. For example, **you can not delete or move externally-managed data**. This is to prevent inconsistencies when these applications want to access their own project data.

* You can **add** [**files**](/project/p-data) and data such as [samples](/project/p-samples) to externally managed projects. Separation of data is ensured by only allowing additional files at the **root level** or in **dedicated subfolders** which you can create in your projects. **Data which you have added can be moved and deleted** again.
* You can **add** [**bundles**](/home/h-bundles) to externally managed projects, provided those bundles do not come with additional restrictions for the project.
* You can **start** [**bench**](/project/p-bench) **workspaces** in externally-managed projects. The resulting data will be stored in the externally-managed project.
* Project administrators and tenant administrators can **disable** [**data sharing**](#create-new-project) on externally managed projects at **Projects > externally\_managed\_project > Project Settings > Details** to prevent data from being copied or extracted.

{% hint style="info" %}
Tertiary modules such as [cohorts](/project/p-cohorts) are not supported for externally-managed projects.
{% endhint %}

Projects are indicated as **externally-managed** in the projects overview screen by a project card with a **managed by app \<app name>** label.

You can keep track of **which files are externally controlled** and which are Platform Core-managed by means of the “**managed by**” column, visible in the data list view of externally-managed projects at **Projects > your\_project > Data**.

### Data Transfer

If you have an externally-manage project and want to **move** the data to another project, you need to:

1. [Copy the data](/project/p-data#data-management) from the externally-managed project to the other (new) project.
2. From within your external application, delete the data which is stored in the externally-managed project in Platform Core.

### Notes

{% hint style="info" %}
When you create a folder with a name which already exists as externally-managed folder, your project will have that folder twice. Once Platform Core-managed and once externally-managed.
{% endhint %}

Externally-managed projects protect their **notification subscriptions** to ensure no user can delete them. It is possible to add your own subscriptions to externally-managed projects, see [notifications](/project/p-notifications) for more information.

## Tutorial

For a better understanding of how all components of Platform Core work, try the [end-to-end tutorial](https://help.ica.illumina.com/tutorials/end-to-end-1).

## Sharing

You can share links to your project and content within projects to people who have [access](/project/p-team) to it. Sharing is done by copying the URL from your browser. This URL contains both the filters and the sort options which you have applied.


# Bundles

Bundles are curated data sets which combine assets such as pipelines, tools, and Base query templates. This is where you will find packaged assets such as Illumina-provided pipelines and sample data. You can create, share and use bundles in projects of your own [tenant](/get-started/gs-getstarted#tenant-setup) as well as projects in other tenants.

{% hint style="info" %}
There is a combined limit of 30 000 projects and bundles per tenant.
{% endhint %}

The following Platform Core assets can be included in bundles:

* [Data](/project/p-data) (link / unlink)
* [Samples](/project/p-samples) (link / unlink)
* [Reference Data ](/project/p-flow/f-referencedata)(add / delete)
* [Pipelines](/project/p-flow/f-pipelines) (link/unlink)
* [Tools](/home/h-toolrepository) and [Tool images](/home/h-dockerrepository) (link/unlink)
* [Base tables](/project/p-base/base-tables) (read-only) (link/unlink)
* [Base query templates](/project/p-base/base-query)
* [Bench docker images](/home/h-dockerrepository)

The main Bundles screen has two tabs: **My Bundles** and **Entitled Bundles**. The **My Bundles** tab shows all the bundles that you are a member of. This tab is where most of your interactions with bundles occur. The **Entitled Bundles** tab shows the bundles that have been specially created by Illumina or other organizations and shared with you to use in your projects. See [Access and Use an Entitled Bundle](https://help.ica.illumina.com/home/h-bundles#access-and-use-an-entitled-bundle).

{% hint style="warning" %}
Some bundles come with additional restrictions such as disabling bench access or internet access when running pipelines to protect the data contained in them. When you link these bundles, the restrictions will be enforced on your project. Unlinking the bundle will not remove the restrictions.

You can not link bundles which come with additional restrictions to [externally managed projects](/home/h-projects#externally-managed-projects).
{% endhint %}

{% hint style="warning" %}
As of Platform Core v.2.29, the content in bundles is linked in such a way that any updates to a bundle are automatically propagated to the projects which have that bundle linked.

If you have created bundle links in Platform Core versions prior to Platform Core v2.29 and want to switch them over to links with dynamic updates, you need to unlink and relink them.
{% endhint %}

## Linking an Existing Bundle to a Project

1. From the main navigation page, select **Projects > your\_project > Project Settings > Details**.
2. Click the **Edit** button at the top of the Details page.
3. Click the **+** button, under **Linked bundles**.
4. Click on the desired bundle, then click the **+Link Bundles** button.
5. Click **Save**.

The assets included in the bundle will now be available in the respective pages within the Project (e.g. Data and Pipelines pages). Any updates to the assets will be automatically available in the destination project.

To **unlink a bundle** from a project,

1. Select **Projects > your\_project > Project Settings > Details**.
2. Click the **Edit** button at the top of the Details page.
3. Click the (**-**) button, next to the linked bundle you wish to remove.

{% hint style="warning" %}
Bundles and projects have to be in the same region in order to be linked. Otherwise, the error *The bundle is in a different region than the project so it's not eligible for linking* will be displayed.
{% endhint %}

{% hint style="info" %}
The owning tenant of a project must have access to a bundle if you want to link that bundle to the project. You do not carry your access to a bundle over if you are invited to projects of other tenants.
{% endhint %}

{% hint style="info" %}
When linking a bundle which includes Base to a project that does not have Base enabled, there are two possibilities:

* **Base is not allowed due to entitlements**: The bundle will be linked and you will be given access to the data, pipelines, samples,... but you will not see the Base tables in your project.
* **Base is allowed, but not yet enabled for the project.** The bundle will be linked and you will be given access to the data, pipelines, samples,...but you will not see the Base tables in your project and Base remains disabled until you enable it.
  {% endhint %}

{% hint style="info" %}
You can not **unlink** bundles which were linked by external applications
{% endhint %}

## Create a New Bundle

To create a new bundle and configure its settings, do as follows.

1. From the main navigation, select **Bundles**.
2. Select **+ Create**.
3. Enter a **unique name** for the bundle.
4. From the Region drop-down list, select where the assets for this bundle should be stored.
5. Set the **status** of the bundle. When the status of a bundle changes, it cannot be reverted to a draft or released state.
   * **Draft**—The bundle can be edited.
   * **Released**—The bundle is released. Technically, you can still edit bundle information and add assets to the bundle, but should refrain from doing so.
   * **Deprecated**—The bundle is no longer intended for use. By default, deprecated bundles are hidden on the main Bundles screen (unless non-deprecated versions of the bundle exist). Select "Show deprecated bundles" to show all deprecated bundles. Bundles can not be recovered from deprecated status.
6. \[optional] Configure the following settings.
   * Categories—Select an existing category or enter a new one.
   * Short Description—Enter a description for the bundle.
   * Metadata Model—Select a metadata model to apply to the bundle.
7. Enter a **release version** for the bundle and optionally enter a description for the version.
8. \[Optional] Links can be added with a display name (max 100 chars) and URL (max 2048 chars).
   * Homepage
   * License
   * Links
   * Publications
9. \[Optional] Enter any information you would like to distribute with the bundle in the Documentation section.
10. Select Save.

{% hint style="warning" %}
There is no option to delete bundles, they must be deprecated instead.
{% endhint %}

{% hint style="info" %}
To **cancel** creating a bundle, select Bundles from the navigation at the top of the screen to return to your bundles overview.
{% endhint %}

## Edit an Existing Bundle

To make changes to a bundle:

1. From the main navigation, select **Bundles**.
2. Select a bundle.
3. Select **Edit**.
4. Modify the bundle information and documentation as needed.
5. Select **Save**.

{% hint style="info" %}
When the changes are saved, they also become available in all projects that have this bundle linked.
{% endhint %}

## Adding Assets to a Bundle

To add assets to a bundle:

1. Select a bundle.
2. On the left-hand side, select the type of asset (such as **Flow > pipelines**, **Base > Tables** or **Bench > Docker Images**) you want to add to the bundle.
3. Select **link** to add assets to the bundle.
4. Select the assets and confirm with the **link** button..

Assets must meet the following requirements before they can be added to a bundle:

* For samples and data, the [project](/home/h-projects) the asset belongs to must have data sharing enabled.
* The region of the project containing the asset must match the region of the bundle.
* You must have permission to access the project containing the asset.
* Pipelines and tools need to be in released status.
* [Samples](/project/p-samples) must be available in a `complete` state.
* [Tables](/project/p-base/base-tables) can not be linked to a bundle while Base creation for that bundle is in progress.

When you link folders to a bundle, a warning is displayed indicating that, depending on the size of the folder, linking may take considerable time. The linking process will run in the background and the progress can be monitored on the **Bundles > your\_bundle > activity > Batch Jobs** screen. To see more details and the progress, double-click the batch job and then double-click the individual item. This will show how many individual files are already linked.

{% hint style="info" %}
You can not add the same asset twice to a bundle. Once added, the asset will no longer appear in the asset selection list.
{% endhint %}

{% hint style="info" %}
You need to be in list view in order to **unlink** items from a bundle. Select the item and choose the unlink action at the top of the screen.
{% endhint %}

Which batch jobs are visible as activity depends on the user role.

## Create a New Bundle Version

When creating a new bundle version, you can only add assets to the bundle. You cannot remove existing assets from a bundle when creating a new version. If you need to remove assets from a bundle, it is recommended that you create a new bundle. All users wich currently have access to a bundle will automatically have access to the new version as well.

1. From the main navigation, select **Bundles**.
2. Select a bundle.
3. Select the + **Create new Version** button.
4. Make updates as needed and **update the version number**.
5. Select **Save**.

When you create a new version of a bundle, it will replace the old version in your list. To see the old version, open your new bundle and look at **Bundles > your\_bundle > Details > Versioning**. There you can open the previous version which is contained in your new version.

Assets such as data which were added in a previous version of your bundle will be marked in green, while new content will be black.

## Add Terms of Use to a Bundle

1. From the main navigation, Select **Bundles > your\_bundle > Bundle Settings > Legal** from the left hand navigation.
2. To add Terms of Use to a Bundle, do as follows:
   * Select **+ Create New Version**.
   * Use the editor to define Terms of Use for the selected bundle.
   * Click Save.
   * \[Optional] Require acceptance by clicking the checkbox next to Acceptance required.\
     `Acceptance required will prompt a user to accept the Terms of Use before being able to use a bundle or add the bundle to a project.`
3. To edit the Terms of Use, repeat Steps 1-3 and use a unique version name. If you select acceptance required, you can choose to keep the acceptance status as is or require users to reaccept the terms of use. When reacceptance is required, users need to reaccept the terms in order continue using this bundle in their pipelines. This is indicated when they want to enter projects which use this bundle.

## Collaborating on a Bundle

If you want to collaborate with other people on creating a bundle and managing the assets in the bundle, you can add users to your bundle and set their permissions. You use this to create a bundle together, not to use the bundle in your projects.

1. From the main navigation, select **Bundles > your\_bundle > Bundle Settings > Team**.
2. To invite a user to collaborate on the bundle, do as follows.

   * To add a user from your tenant, select **Someone of your tenant** and select a user from the drop-down list.
   * To add a user by their email address, select **By email** and enter their email address.
   * To add all the users of an entire workgroup, select **Add workgroup** and select a workgroup from the drop-down list.
   * Select the Bundle Role drop-down list and choose a role for the user or workgroup. This role defines the ability of the user or workgroup to view or edit bundle settings.
     * **Viewer**: view content without editing rights.
     * **Contributor**: view bundle content and link/unlink assets.
     * **Administrator**: full edit rights of content and configuration.
   * Repeat as needed to add more users.

   Users are not officially added to the bundle until they accept the invitation.
3. To change the permissions role for a user, select the Bundle Role drop-down list for the user and select a new role.
4. To revoke bundle permissions from a user, select the trash icon for the user.
5. Select **Save Changes**.

## Sharing a Bundle

Once you have finalized your bundle and added all assets and legal requirements, you can share your bundle with other tenants to use it in their projects.

Your bundle must be in released status to prevent it from being updated while it is shared.

1. Go to **Bundles > your\_bundle > Edit > Details > Bundle status** and set it to Released.
2. Save the change.

Once the bundle is released, you can share it. Invitations are sent to an individual email address, however **access is granted and extended to all users and all workgroups inside that tenant**.

1. Go to **Bundles > your\_bundle > Bundle Settings > Share**.
2. Click **Invite** and enter the email address of the person you want to share the bundle with. They will receive an email from which they can accept or reject the invitation to use the bundle. The invitation will show the bundle name, description and owner. The link in the invite can only be used once.

{% hint style="info" %}
Do not to create duplicate entries. You can only use one user/tenant combination per bundle.
{% endhint %}

You can follow up on the status of the invitation on the **Bundles > your\_bundle > Bundle Settings > Share** page.

* If they **reject** the bundle, the rejection date will be shown.\
  To re-invite that person again later on, **select their email address** in the list and choose **Remove**. You can then create a **new invitation**. If you do not remove the old entry before sending a new invitation, they will be unable to accept and get an error message stating that the user and bundle combination must be unique. They can also not re-use an invitation once it has been accepted or declined.
* If they **accept** the bundle, the acceptance date will be shown. They will in turn see the bundle under **Bundles > Entitled bundles**.\
  To **remove** access, **select their email address** in the list and choose Remove.

## Entitled Bundles

Entitled bundles are bundles created by Illumina or third parties for you to use in your projects. Entitled bundles can already be part of your tenant when it is part of your subscription. You can see your entitled bundles at **Bundles > Entitled Bundles**.

To use your shared entitled bundle, add the bundle to your project via Project Linking.

1. From the main navigation page, select **Projects > your\_project > Project Settings > Details**.
2. Click the **Edit** button at the top of the Details page.
3. Click the **+** button, under **Linked bundles**.
4. Click on the desired bundle, then click the **+Link Bundles** button.
5. Click **Save**.

Content shared via entitled bundles is read-only, so you cannot add or modify the contents of an entitled bundle. If you lose access to an entitled bundle previously shared with you, the bundle is unlinked and you will no longer be able to access its contents.


# Event Log

The event log shows an overview of system events with options to search and filter. For every entry, it lists the following:

* Event date and time
* Category (error, warn or info)
* Code
* Description
* Tenant

Up to 200,000 results will be be returned. If your desired records are outside the range of the returned records, please refine the filters or use the **search** function at the top right.

**Export** is restricted to the amount of entries shown per page. You can use the selector at the bottom to set this to up to 1000 entries per page.


# Metadata Models

Illumina Connected Analytics allows you to create and assign metadata to **capture additional information about samples**.

Every tenant has a **root metadata model** that is accessible to all projects of that tenant. This allows an organization to collect the same piece of information, such as an ID number, for every sample in every project. Within this root model, you can configure **multiple metadata submodels,** even at different levels. These submodels inherit all fields and groups from their parent models.

<figure><img src="https://documents.lucid.app/documents/32851362-0671-4656-b0a1-478fdc25c44b/pages/0_0?a=949&#x26;x=1983&#x26;y=1016&#x26;w=860&#x26;h=494&#x26;store=1&#x26;accept=image%2F*&#x26;auth=LCA%20635a1cd5870675172bee1dbb2e4ce43600ae4702e825f4fa89f04527e2cccf09-ts%3D1754556661" alt="" width="375"><figcaption></figcaption></figure>

Illumina recommends that you **limit the amount of fields or field groups you add to the root model**. **Fields** can have various types containing single or multiple values and **field groups** contain fields that belong together, such as all fields related to quality metrics. If there are any misconfigured items in the root model, it will carry over into all other tenant metadata models. **Once a root model is published, the fields and groups that are defined within it cannot be deleted, only more fields can be added**.

{% hint style="info" %}
Illumina recommends that you limit the amount of fields or field groups you add to the root model as this model can not be deprecated and anything you add to the root model can not be removed. You should always consider creating submodels before adding anything to the root model.
{% endhint %}

{% hint style="warning" %}
Do not use dots (.) in the metadata model names, fieldgroup names or field names as this can cause issues with field data.
{% endhint %}

When configuring a project, you can assign a published metadata model for all samples in the project. This metadata model can be any published metadata model in your tenant such as the root model, or one of the lower level submodels. When a metadata model is selected for a project, all fields configured for the metadata model, and all fields in any parent models are applied to the samples in the project.

Metadata gives information about a sample and can be provided by the user, the pipeline and the API. There are 2 general categories of metadata models: **Project Metadata** models and **Pipeline Metadata** models . Both models contain metadata fields and groups.

* The **project** metadata model is **specific per tenant.** A Project metadata model has metadata linked to a specific project. **Values are known upfront**, general information is required for each sample of a specific project, and it may include general mandatory company information.
* The **pipeline** metadata model is **linked to a pipeline**, not to a project and can be **shared across tenants**. Values are **populated during pipeline execution** and it requires an output file with the name 'metadata.response.json'.

{% hint style="info" %}
Field groups should be used when configuring metadata fields that are filled by a pipeline. These fields should be part of the same field group and be configured with the Multiple Value setting enabled.
{% endhint %}

Each sample can have multiple metadata models. When you link a project metadata model to your project, you will see its groups and fields present on each sample. The root model from that tenant will be present as every metadata model inherits the groups and fields specified in the parent metadata model(s). When a pipeline is executed with single sample and the pipeline containing a metadata model, the groups and fields will be present as well for each analysis resulting from a pipeline execution.

## Creating a Metadata Model

In the main navigation, go to **System Settings > Metadata Models**. Here you will see the root metadata model and any underlying sub-metadata models. To create a new submodel, select **+Create** at the bottom of the screen.

The new metadata model screen will show an overview of all the higher-level metadata models. use the down arrow next to the model name to expand these for more information.\
![](/files/jDmMVhruFWguQRLfzIlT)

For your new metadata model, add a unique **name** and optional **description**. Once this is done, start adding the metadata fields with the **+Add** button. The field type will determine the parameters which you can configure.

To edit your metadata model later on, select it and choose **Manage > Edit**. Keep in mind that fields can be added, but not removed once the model is published.

### Field Types & Properties

<table><thead><tr><th width="180.85546875">field types</th><th></th></tr></thead><tbody><tr><td>Text</td><td>Free text</td></tr><tr><td>Keyword</td><td>Automatically complete value based on already used values</td></tr><tr><td>Numeric</td><td>Only numbers</td></tr><tr><td>Boolean</td><td>True or false, cannot be multiple value</td></tr><tr><td>Date</td><td>e.g. 2021-09-20</td></tr><tr><td>Date time</td><td>e.g. 2021-09-20 11:43:53, saved in UTC</td></tr><tr><td>Enumeration</td><td>select value from list. <em>Enter the values in the options field which appears when you have selected enumeration type.</em></td></tr><tr><td>Field Group</td><td>Groups fields. Once you have chosen this, the +Add group field becomes available to add fields to this group.</td></tr></tbody></table>

<figure><img src="/files/MNZaPsETxEROoNlr2HW4" alt=""><figcaption></figcaption></figure>

The following properties can be selected for groups & fields:

<table><thead><tr><th width="171.19140625">Propery</th><th></th></tr></thead><tbody><tr><td>Required</td><td>Pipeline can not be started with this sample until the required group/field is filled in.</td></tr><tr><td>Sensitive</td><td>Values of this group/field are only visible to project users of the own tenant. When a sample is shared across tenants, these fields will not be visible.</td></tr><tr><td>Multi value</td><td>This group/field can consist of multiple (grouped) values</td></tr><tr><td>Filled by pipeline</td><td><p>Fields that need to be filled by pipeline should be part of the same group. This group will automatically be multiple value and values will be available after pipeline execution. <em>This property is only available for the Field Group type.</em></p><p>If you have fields that are filled by the pipeline you can create an <strong>example JSON</strong> structure indicating what the json in an analysis output file with name metadata.response.json should look like to fill in the metadata fields of this model. Use <strong>System Settings > Metadata Models > your_metadata_model > Manage > Generate example JSON</strong>. Only fields in groups marked as <em>Filled by pipeline</em> are included.</p></td></tr></tbody></table>

{% hint style="warning" %}
Fields cannot be both **required** and **filled by pipeline** at the same time.
{% endhint %}

{% hint style="info" %}
To help retrieve the field values via API calls, you can use **System Settings > Metadata Models > your\_metadata\_model > Manage > Show Field Paths**.
{% endhint %}

## Metadata Actions

### Publish a Metadata Model

Newly created and updated metadata models are not available for use within the tenant until the metadata model is published. Once a metadata model is published, fields and field groups cannot be deleted, but the names and descriptions for fields and field groups can be edited. A model can be published after verifying all parent models are published first. To publish your model, select **System Settings > Metadata Models > your\_metadata\_model > Manage > Publish**.

### Retire a Metadata Model

If a published metadata model is no longer needed, you can retire the model (except the root model). Once a model is retired, it can be published again in case you would need to reactivate it.

1. First, check if the model contains any submodels. A model cannot be retired if it contains any published submodels.
2. When you are certain you want to retire a model and all submodels are retired, select **System Settings > Metadata Models > your\_metadata\_model > Manage > Retire Metadata Model**.

### Assign a Metadata Model to a Project

To add metadata to your samples, you first need to assign a metadata model to your project.

1. Go to **Projects > your\_project > Project Settings > Details**.
2. Select **Edit**.
3. From the **Metadata Model** drop-down list, select the metadata model you want to use for the project.
4. Select **Save**. All fields configured for the metadata model, and all fields in any parent models are applied to the samples in the project.

### Add Metadata to Samples Manually

If you have a metadata model assigned to your project, you can manually fill out the defined metadata of the samples in your project:

1. Go to **Projects > your\_project > Samples > your\_sample**.
2. Click your sample to open the sample details and choose **Edit Sample**.
3. Enter all metadata information as it applies to the selected sample. All required metadata fields must be populated or the pipeline will not be able to start.
4. Select **Save**

### Populating a Pipeline Metadata Model

To fill metadata by pipeline executions, a pipeline model must be created.

1. In the main navigation, go to **Projects > your\_project > Flow > Pipelines > your\_pipeline**.
2. Click on your pipeline to open the pipeline details and choose **Edit**.
3. Create/Edit your model under **Metadata Model tab**. Field groups should be used when configuring metadata fields that are filled by a pipeline. These fields should be part of the same field group and be configured with the Multiple Value setting enabled.

In order for your pipeline to fill the metadata model, an output file with the name `metadata.response.json` must be generated. After adding your group fields to the pipeline model, click on `Generate example JSON` to view the required format for your pipeline.

Use **System Settings > Metadata Models > your\_metadata\_model > Manage > Generate example JSON** to see an example JSON for these fields.

{% hint style="warning" %}
The field names cannot have `.` in them, e.g. for the metric name `Q30 bases (excl. dup & clipped bases)` the `.` after `excl` must be removed.
{% endhint %}

## Pushing Metadata Metrics to Base

Populating metadata models of samples allows having a sample-centric view of all the metadata. It is also possible to synchronize that data into your project's Base warehouse.

1. In Platform Core, select **Projects > your\_project >Base > Schedule**.
2. Select **+Create > From metadata**.
3. Type a name for your schedule, optionally add a description, and set it to active. You can select if sensitive metadata fields should be included as values of sensitive metadata fields will not be visible to other users outside of the project.
4. Select Save.
5. Navigate to **Base > Tables** in your project.
6. Two new table schemas should be added with your current metadata models.


# Docker Repository

In order to create a Tool or Bench image, a Docker image is required to run the application in a containerized environment.\
Illumina BioInsight Platform Core supports both public Docker images and private Docker images uploaded to Platform Core.

{% hint style="warning" %}
Use Docker images built for x86 architecture or multi-platform images that support x86. You can build Docker images that support [both ARM and X86](https://docs.docker.com/build/building/multi-platform/) structure.
{% endhint %}

## Importing a Public External Image (Tools)

1. Navigate to **System Settings > Docker Repository**.
2. Click **Create > External image** to add a new external image.
3. Add your full image URL in the Url field, e.g. `docker.io/alpine:latest` or `registry.hub.docker.com/library/alpine:latest`. Docker Name and Version will auto-populate. (Tip: do not add http\:// or https\:// in your URL)

{% hint style="warning" %}
Do not use **:latest** when the repository has rate limiting enabled as this interferes with caching and incurs additional data transfer.
{% endhint %}

4. (Optional) Complete the Description field.
5. Click **Save**.
6. The newly added image will appear in your Docker Repository list. You can differentiate between internal and external images by looking at the **Source** column. If this column is not visible, you can add it with the columns icon (<img src="/files/WMMgm1uYUEW0180V5Rz1" alt="" data-size="line">).

{% hint style="info" %}
Verification of the URL is performed during execution of a pipeline which depends on the Docker image, not during configuration.
{% endhint %}

{% hint style="info" %}
External images are accessed from the external source whenever required and not stored in Platform Core. Therefore, it is important **not to move or delete the external source**. There is no status displayed on external Docker repositories in the overview as Platform Core cannot guarantee their availability.

The use of **:stable** instead of :latest is recommended.
{% endhint %}

## Importing a Private Image (Tools + Bench Images)

In order to use private images in your tool, you must first upload them as a TAR file.

1. Navigate to **Projects > your\_project** .
2. **Upload your private image as a TAR file**, either by dragging and dropping the file in the Data tab, using the CLI or a Connector. For more information please refer to project [Data](/project/p-data#uploaddata).
3. Select your uploaded TAR file and click in the top menu on **Manage > Change Format** .
4. Select **DOCKER** from the drop-down menu and Save.

   <figure><img src="/files/zj5paVuFawoWXkCiKJXI" alt=""><figcaption></figcaption></figure>
5. Navigate to **System Settings > Docker Repository** (outside of your project).
6. Click on **Create > Image**.
7. Click on the Docker image field. Select the TAR file from the desired region.

<figure><img src="/files/ZOKUQpwu0u6qxgd8D0Ch" alt=""><figcaption></figcaption></figure>

8. Provide a **name** and **version** for your Docker image, this will automatically populate the global URL field.
9. Select the appropriate region, determine the type (tool or bench image), the [cluster](/project/p-bench/bench-workspaces) compatibility (only available for bench images), access method and click Save.
10. The newly added image should appear in your Docker Repository list. Verify it is marked as Available under the Status column to ensure it is ready to be used in your tool or pipeline.

## Copying Docker Images to other Regions

1. Navigate to **System Settings > Docker Repository**.
2. Either
   * Select the required image(s) and go to **Manage > Add Region**. Select the desired active region and chose **Add**.
   * OR open the image details, check the box matching the active region you want to add, and select **save**.
3. In both cases, allow a few minutes for the image to become available in the new region (the status becomes available in table view).

To **remove regions**, go to **Manage > Remove Region** or unselect the regions from the Docker image detail view.

## Downloading Docker Images

You can download your created Docker images at **System Settings > Docker Images > your\_Docker\_image > Manage > Download**.

In order to be able to download Docker images, the following requirements must be met:

* The Docker image can **not** be from an **entitled** bundle.
* Only **self-created** Docker images can be downloaded.
* The Docker image must be an **internal** image and in status **Available**.
* You can only select a single Docker image at a time for download.
* You need a [**service** **connector**](/project/p-connectivity/service-connector) with a download rule to download the Docker image.

## Deleting Docker Images

When you no longer need a specific Docker image, go to **System Settings > Docker Images > your\_Docker\_image > Manage > Delete**. A check is performed to see if this Docker image is still used in any workspaces. Depending on your user rights, you may not get the delete option or be presented with a warning indicating in how many workspaces the Docker image is still being used Proceeding with the deletion will invalidate the workspaces which use this image.

Docker images are not physically deleted, but get marked as deleted. You can still see them if you select Show deleted Docker images in the **System Settings > Docker Images** view.

## File Size Considerations

Docker image size should be kept as small as practically possible. To this end, it is best practice to compress the image. After compressing and uploading the image, select your uploaded file and click **Manage > Change Format** in the top menu to change it to Docker format so Platform Core can recognize the file.


# Tool Repository

A Tool is the definition of a containerized application with defined inputs, outputs, and execution environment details including compute resources required, environment variables, command line arguments, and more.

## Create a Tool

Tools define the inputs, parameters, and outputs for the analysis. Tools are available for use in graphical Common Workflow Language (CWL) pipelines by any project in the account.

1. Select **System Settings > Tool Repository > + Create**.
2. Configure tool settings in the tool properties tabs. See [Tool Properties](#tool-properties).
3. Select Save.

The following sections describe the tool properties that can be configured in each tab.

{% hint style="info" %}
Refer to the [CWL CommandLineTool Specification](https://www.commonwl.org/v1.0/CommandLineTool.html) for further explanation about many of the properties described below. Not all features described in the specification are supported.
{% endhint %}

### Details Tab

<table><thead><tr><th width="214">Field</th><th>Entry</th></tr></thead><tbody><tr><td>Name</td><td>The name of the tool.</td></tr><tr><td>Description</td><td>Free text description for information purposes.</td></tr><tr><td>Icon</td><td>The icon for the tool.</td></tr><tr><td>Status</td><td>The release <a href="#tool-status">status</a> of the tool.</td></tr><tr><td>Docker image</td><td>The registered Docker image for the tool.</td></tr><tr><td>Categories</td><td>One or more tags to categorize the tool. Select from existing tags or type a new tag name in the field.</td></tr><tr><td>Tool version</td><td>The version of the tool specified by the end user. Could be any string.</td></tr><tr><td>Release version</td><td>The version number of the tool.</td></tr><tr><td>Version comment</td><td>A description of changes in the updated version.</td></tr><tr><td>Links</td><td>External reference links.</td></tr><tr><td>Documentation</td><td>The Documentation field provides options for configuring the HTML description for the tool. The description appears in the Tool Repository but is excluded from exported CWL definitions.</td></tr></tbody></table>

#### **Tool Status**

The tool release status can be set to *Draft*, *Release Candidate*, *Released* or *Deprecated*.

The *Building* and *Build Failed* options are set by the application and not during configuration.

<table><thead><tr><th width="190">Status</th><th>Description</th></tr></thead><tbody><tr><td>Draft</td><td>Fully editable draft.</td></tr><tr><td>Release Candidate</td><td>The tool is ready for release. Editing is locked but the tool can be cloned to create a new version.</td></tr><tr><td>Released</td><td>The tool is released. Tools in this state cannot be edited. Editing is locked but the tool can be cloned to create a new version.</td></tr><tr><td>Deprecated</td><td>The tool is no longer intended for use in pipelines. but there are no restrictions placed on the tool. That is, it can still be added to new pipelines and will continue to work in existing pipelines. It is merely an indication to the user that the tool should no longer be used.</td></tr></tbody></table>

#### General Tab

The General provides options to configure the basic command line.

<table><thead><tr><th width="194">Field</th><th>Entry</th></tr></thead><tbody><tr><td>ID</td><td>CWL identifier field</td></tr><tr><td>CWL version</td><td>The CWL version in use. This field cannot be changed.</td></tr><tr><td>Base command</td><td>Components of the command. Each argument must be added in a separate line.</td></tr><tr><td>Standard out</td><td>The name of the file where the Standard Out (STDOUT) stream information will be stored.</td></tr><tr><td>Standard error</td><td>The name of the file where the Standard Error (STDERR) stream information will be stored.</td></tr><tr><td>Requirements</td><td>The requirements for triggering an error message. (see below)</td></tr><tr><td>Hints</td><td>The requirements for triggering a warning message. (see below)</td></tr></tbody></table>

The Hints/Requirements include CWL features to indicate capabilities expected in the Tool's execution environment.

* **Inline Javascript**
  * The Tool contains a property with a JavaScript expression to resolve it's value.
* **Initial workdir**
  * The workdir can be any of the following types:
    * String or Expression — A string or JavaScript expression, eg, `$(inputs.InputFASTA)`
    * File or Dir — A map of one or more files or directories, in the following format: `{type: array, items: [File, Directory]}`
    * Dirent — A script in the working directory. The Entry name field specifies the file name.
* **Scatter feature** — Indicates that the workflow platform must support the `scatter` and `scatterMethod` fields.

#### Arguments Tab

The Arguments tab provides options to configure base command parameters that do not require user input.

Tool arguments may be one of two types:

* **String or Expression** — A literal string or JavaScript expression, eg --format=bam.
* **Binding** — An argument constructed from the binding of an input parameter.

The following table describes the argument input fields.

<table><thead><tr><th width="156">Field</th><th width="388">Entry</th><th>Type</th></tr></thead><tbody><tr><td>Value</td><td>The literal string to be added to the base command.</td><td>String or expression</td></tr><tr><td>Position</td><td>The position of the argument in the final command line. If the position is not specified, the default value is set to 0 and the arguments appear in the order they were added.</td><td>Binding</td></tr><tr><td>Prefix</td><td>The string prefix.</td><td>Binding</td></tr><tr><td>Item separator</td><td>The separator that is used between array values.</td><td>Binding</td></tr><tr><td>Value from</td><td>The source string or JavaScript expression.</td><td>Binding</td></tr><tr><td>Separate</td><td>The setting to require the Prefix and Value from fields to be added as separate or combined arguments. Tru indicates the fields must be added as separate arguments. False indicates the fields must be added as a single concatenated argument.</td><td>Binding</td></tr><tr><td>Shell quote</td><td>The setting to quote the Value from field on the command line. True indicates the value field appears in the command line. False indicates the value field is entered manually.</td><td>Binding</td></tr></tbody></table>

**Example**

<table><thead><tr><th width="169">Field</th><th>Value</th></tr></thead><tbody><tr><td>Prefix</td><td><code>--output-filename</code></td></tr><tr><td>Value from</td><td><code>$(inputs.inputSAM.nameroot).bam</code></td></tr><tr><td>Input file</td><td><code>/tmp/storage/SRR45678_sorted.sam</code></td></tr><tr><td>Output file</td><td><code>SRR45678_sorted.bam</code></td></tr></tbody></table>

#### Inputs Tab

The Inputs tab provides options to define the input files and folders for the tool. The following table describes the input and binding fields. Selecting multi value enables type binding options for adding prefixes to the input.

<table><thead><tr><th width="184">Field</th><th>Entry</th></tr></thead><tbody><tr><td>ID</td><td>The file ID.</td></tr><tr><td>Label</td><td>A short description of the input.</td></tr><tr><td>Description</td><td>A long description of the input.</td></tr><tr><td>Type</td><td>The input type, which can be either a file or a directory.</td></tr><tr><td>Input options</td><td><strong>Optional</strong> indicates the input is optional.<br><strong>Multi value</strong> indicates there is more than one input file or directory.<br><strong>Streamable</strong> indicates the file is read or written sequentially without seeking.</td></tr><tr><td>Secondary files</td><td>The required secondary files or directories.</td></tr><tr><td>Format</td><td>The input file format.</td></tr><tr><td>Position</td><td>The position of the argument in the final command line. If the position is not specified, the default value is set to 0 and the arguments appear in the order they were added.</td></tr><tr><td>Prefix</td><td>The string prefix.</td></tr><tr><td>Item separator</td><td>The separator that is used between array values.</td></tr><tr><td>Value from</td><td>The source string or JavaScript expression.</td></tr><tr><td>Load contents</td><td>The setting to require the Prefix and Value from fields to be added as separate or combined arguments. True indicates the fields must be added as separate arguments. False indicates the fields must be added as a single concatenated argument.</td></tr><tr><td>Separate</td><td>The setting to require the Prefix and Value from fields to be added as separate or combined arguments. True indicates the fields must be added as separate arguments. False indicates the fields must be added as a single concatenated argument.</td></tr><tr><td>Shell quote</td><td>The setting to quote the Value from field on the command line. True indicates the value field appears in the command line. False indicates the value field is entered manually.</td></tr></tbody></table>

#### Settings Tab

The Settings tab provides options to define parameters that can be set at the time of execution. The following table describes the input and binding fields. Selecting multi value enables type binding options for adding prefixes to the input.

<table><thead><tr><th width="205">Field</th><th>Entry</th></tr></thead><tbody><tr><td>ID</td><td>The file ID.</td></tr><tr><td>Label</td><td>A short description of the input.</td></tr><tr><td>Description</td><td>A long description of the input.</td></tr><tr><td>Type</td><td>The input type, which can be Boolean, Int, Long, Float, Double or String.</td></tr><tr><td>Default Value</td><td>The default value to use if the tool setting is not available.</td></tr><tr><td>Input options</td><td><strong>Optional</strong> indicates the input is optional.<br><strong>Multi value</strong> indicates there can be more than one value for the input.</td></tr><tr><td>Position</td><td>The position of the argument in the final command line. If the position is not specified, the default value is set to 0 and the arguments appear in the order they were added.</td></tr><tr><td>Prefix</td><td>The string prefix.</td></tr><tr><td>Item separator</td><td>The separator that is used between array values.</td></tr><tr><td>Value from</td><td>The source string or JavaScript expression.</td></tr><tr><td>Separate</td><td>The setting to require the Prefix and Value from fields to be added as separate or combined arguments. True indicates the fields must be added as separate arguments. False indicates the fields must be added as a single concatenated argument.</td></tr><tr><td>Shell quote</td><td>The setting to quote the Value from field on the command line. True indicates the value field appears in the command line. False indicates the value field is entered manually.</td></tr></tbody></table>

#### Outputs Tab

The Outputs tab provides options to define the parameters of output files.

The following table describes the input and binding fields. Selecting multi value enables type binding options for adding prefixes to the input.

<table><thead><tr><th width="174">Field</th><th>Entry</th></tr></thead><tbody><tr><td>ID</td><td>The file ID.</td></tr><tr><td>Label</td><td>A short description of the input.</td></tr><tr><td>Description</td><td>A long description of the input.</td></tr><tr><td>Type</td><td>The input type, which can be either a file or a directory.</td></tr><tr><td>Output options</td><td><strong>Optional</strong> indicates the input is optional.<br><strong>Multi value</strong> indicates here is more than one input file or directory.<br><strong>Streamable</strong> indicates the file is read or written sequentially without seeking.</td></tr><tr><td>Secondary files</td><td>The required secondary files or folders.</td></tr><tr><td>Format</td><td>The input file format.</td></tr><tr><td>Globs</td><td>The pattern for searching file names.</td></tr><tr><td>Load contents</td><td>Automatically loads some contents. The system extracts up to the first 64 KiB of text from the file. Populates the contents field with the first 64 KiB of text from the file.</td></tr><tr><td>Output eval</td><td>Evaluate an expression to generate the output value.</td></tr></tbody></table>

## Edit a Tool

As long as your tool is still in draft mode, you can edit it. Once released, you need to clone it to have an editable copy.

1. From the **System Settings > Tool Repository** page, select a tool.
2. Select **Edit**.

### Update Tool Status

1. From the **System Settings > Tool Repository** page, select a tool.
2. Select the **Information** tab.
3. From the Status drop-down menu, select a status.
4. Select Save.

## Creating definitions without the wizard

In addition to the interactive Tool builder, the platform GUI also supports working directly with the raw definition on the right hand side of the screen when developing a new Tool. This provides the ability to write the Tool definition manually or bring an existing Tool's definition to the platform.

{% hint style="warning" %}
Be careful when editing the raw tool definition as this can introduce errors.
{% endhint %}

A simple example CWL Tool definition is provided below.

```yaml
#!/usr/bin/env cwl-runner

cwlVersion: v1.0
class: CommandLineTool
label: echo
inputs:
  message:
    type: string
    default: testMessage
    inputBinding:
      position: 1
outputs:
  echoout:
    type: stdout
baseCommand:
- echo
```

After pasting into the editor, the definition is parsed and the other tabs for visually editing the Tool will populate according to the definition contents.

## Creating Your First Tool - Tips and Tricks

* General Tool - includes your base command and various optional configurations.
  * The base command is required for your tool to run, e.g. `python /path/to/script.py` such that `python` and `/path/to/script.py` are added in separate lines.
  * Inline Javascript requirement - must be enabled if you are using Javascript anywhere in your tool definition.
  * Initial workdir requirement - Dirent Type
    * Your tool must point to a script that executes your analysis. That script can either be provided in your Docker image or using a Dirent. Defining a script via Dirent allows you to dynamically modify your script without updating your Docker image. In order to define your Dirent script define your script name under `Entry name` (e.g. `runner.sh`) and the script content under `Entry`. Then, point your base command to that custom script, e.g. `bash runner.sh`.

{% hint style="info" %}
The difference between **Settings** and **Arguments:** Settings are exposed at the pipeline level with the ability to get modified at launch, while Arguments are intended to be immutable and hidden from users launching the pipeline.
{% endhint %}

* How to reference your tool inputs and settings throughout the tool definition?
  * You can either reference your inputs using their position or ID.
    * Settings can be referenced using their defined IDs, e.g. `$(inputs.InputSetting)`
    * File/Folder inputs can be referenced using their defined IDs, followed by the desired field, e.g. `$(inputs.InputFile.path)`. For additional information please refer to the [File CWL documentation](https://www.commonwl.org/v1.0/CommandLineTool.html#File).
    * All inputs can also be referenced using their position, e.g. `bash script.sh $1 $2`


# Storage

A storage configuration provides Platform Core with information to connect to an external cloud storage provider, such as AWS S3. The storage configuration validates that the information provided is correct, and then continuously monitors the integration.

Refer to the following pages for instructions to setup supported external cloud storage providers:

* [Connect AWS S3 Bucket](/home/h-storage/s-awss3)

## Credentials

The storage configuration requires credentials to connect to your storage. AWS uses the security credentials to authenticate and authorize your requests. On the **System Settings > Credentials > Create > Storage Credential**, you can enter these credentials.

Fill out the following fields:

* **Type** - The type of access credentials. This can be either AWS USER or AWS ROLE.
* **Name** - Provide a name to easily identify your access key.
* **Access key ID (AWS user only)** - The access key you created.
* **Secret access key (AWS user only)** - Your related secret access key.
* **RoleSessionName prefix (AWS role only)** - generated by using the generate button. This value must be included in the [trust policy](/home/h-storage/s-awss3/iam-role-method#trust-policy). for the specified AWS role. **You can only download or copy this value now during creation.** Once this dialog box closes after saving, you can no longer access this value.

You can **share the credentials** you own with other users of your tenant. To do so select your credentials at **System Settings > Credentials** and choose **Share**.

For more information, refer to the [AWS security credentials](https://docs.aws.amazon.com/IAM/latest/UserGuide/security-creds.html) documentation.

## Create a Storage Configuration

1. In the Platform Core main navigation, select **System Settings > Storage > Create**.
2. Configure the following settings for the storage configuration.
   * **Type** - Always use the default value here (AWS\_S3).
   * **Region** - Select the region where the bucket is located.
   * **Configuration name** - You will use this name when creating volumes that reside in the bucket. The name length must be in between 3 and 63 characters.
   * **Description** - Here you can provide a description for yourself or other users to identify this storage configuration.
   * **Bucket name** - Enter the name of your S3 bucket.
   * **Key prefix** - You can provide a key prefix to allow only files inside the prefix to be accessible. Although this setting is optional, **it is highly recommended** to use a key prefix and **mandatory when using** [**dedicated folders in your S3 storage**](/home/h-storage/s-awss3#id-2-create-data-access-permission-aws-iam-policy). **The key prefix must end with "/".**
     * If a **key prefix** is specified, your projects will **only have access to that folder** and subfolders. For example, using the key prefix **folder-1/** ensures that only the data from the **folder-1** folder in your S3 bucket is synced with your Platform Core project. Using prefixes and distinct folders for each Platform Core project is the recommended configuration as it allows you to use the same S3 bucket for different projects.
     * Using **no key prefix** (**not recommended**) results in syncing all data in your S3 bucket (starting from root level) with your Platform Core project. Your project will have access to your entire S3 bucket, which **prevents that S3 bucket from being used for other Platform Core projects**.
   * **Storage Credential** - Select the credentials to associate with this storage configuration. These were created on **System Settings > Credentials > Create > Storage Credential**.
   * **Role ARN (AWS ROLE)** - This value needs to be copied from your [Role](/home/h-storage/s-awss3/iam-role-method#id-4-create-aws-iam-role) in AWS, so this step can only be completed after performing the steps described in the [IAM Role](/home/h-storage/s-awss3/iam-role-method) method.
   * **Server Side Encryption** \[Optional]—If needed, you can enter the **algorithm** and **key name** for server-side encryption processes.
3. Select Save.

Platform Core performs a series of steps in the background to verify the connection to your bucket. This can take several minutes. You may need to manually refresh the list to verify that the bucket was successfully configured. Once the storage configuration setup is complete, the configuration can be used while [creating a new project](/home/h-projects#create-with-storage-configuration).

With the action **Manage > Set as default for region**, you select which storage will be used as default storage in a region for new projects of your tenant. Only one storage can be default at a time for a region, so selecting a new storage as default will unselect the previous default. If you do not want to have a default, you can select the default storage and the action will become **Unset as default for region**.

The **System** **Settings > Credentials > select your credentials > Manage > Share** action is used to make the storage available to everyone in your tenant. By default, storage is private per user so that you have complete control over the contents. Once you decide you want to share the storage, simply select it and use the **Share** action. Do take into account that once shared, you can not unshare the storage. Once your shared storage is used in a project, it can also no longer be deleted.

{% hint style="warning" %}
Filenames beginning with / are not allowed, so be careful when entering full path names. Otherwise the file will end up on S3 but not be visible in Platform Core. If this happens, access your S3 storage directly and copy the data to where it was intended. If you are using an Illumina-managed S3 storage, submit a support request to delete the erroneous data.
{% endhint %}

## Deleting Storage Configurations

In the Platform Core main navigation, select **System Settings > Storage > select your storage > Manage > Delete**. You can then create a new storage configuration to reuse the bucket name and key prefix.

{% hint style="info" %}
[Hiding a project](/home/h-projects#managing-projects) will also unlock your storage configuration so that it can be reused for another project. Data stored by the hidden project will remain in your S3 storage, so you may need to perform manual cleanup before reusing the storage.
{% endhint %}

## Storage Configuration Verification

Every 4 hours, Platform Core will verify the storage configuration and credentials to ensure availability. When an error is detected, Platform Core will attempt to reconnect once every 15 minutes. After 200 consecutively failed connection attempts (50 hours), Platform Core will stop trying to connect.

When you update your credentials, the storage configuration is automatically validated. In addition, you can **manually trigger revalidation** when Platform Core has stopped trying to connect by selecting the storage and then clicking **Validate** on the **System Settings > Storage > select your storage > Manage > Validate**.

Refer to this [page](https://help.ica.illumina.com/reference/r-troubleshootingvolumeconfiguration) for the troubleshooting guide.

## Supported Storage Classes

Platform Core supports the following storage classes. Please see the [AWS documentation](https://aws.amazon.com/s3/storage-classes/) for more information on each:

| Object Class                         | ICA Status |
| ------------------------------------ | ---------- |
| S3 Standard                          | Available  |
| S3 Intelligent-Tiering               | Available  |
| S3 Express One Zone                  | Available  |
| S3 Standard-IA                       | Available  |
| S3 One Zone-IA                       | Available  |
| S3 Glacier Instant Retrieval         | Available  |
| S3 Glacier Flexible Retrieval        | Archived   |
| S3 Glacier Deep Archive              | Archived   |
| Reduced redundancy (not recommended) | Available  |

{% hint style="warning" %}
If you are using [Intelligent Tiering](https://aws.amazon.com/s3/storage-classes/intelligent-tiering/), which allows S3 to automatically move files into different cost-effective storage tiers, please do NOT include the Archive and Deep Archive Access tiers, as these are not supported by Platform Core yet. Instead, you can use lifecycle rules to automatically move files to Archive after 90 days and Deep Archive after 180 days. Lifecycle rules are supported for user-managed buckets.
{% endhint %}


# Connect AWS S3 Bucket

You can use your **own S3 bucket** (unversioned, versioned, versioning-suspended) with Illumina BioInsight Platform Core for data storage. This section describes how to configure your AWS account to allow Platform Core to connect to an S3 bucket.

{% embed url="<https://www.youtube.com/watch?v=5h18XqTgXts&list=PLKRu7cmBQlaiQT6Giou9aSkZ4C0LMIGbc&index=11>" %}
Connect AWS S3 Bucket to Platform Core Project
{% endembed %}

## Prerequisite

#### AWS CLI

These instructions utilize the AWS CLI. Follow the [AWS CLI documentation](https://aws.amazon.com/cli/) for instructions to download and install.

## Best Practices

#### Do not use the root folder of your S3 storage

{% hint style="warning" %}
When configuring a new project in Platform Core to use a preconfigured S3 bucket, **create a folder on your S3 bucket** in the AWS console. This folder will be connected to Platform Core as a prefix.

Failure to create a folder will result in the root folder of your S3 bucket being assigned which will block your S3 bucket from being used for other Platform Core projects with the error "Conflict while updating file/folder. Please try again later."
{% endhint %}

#### Service Control Policies & Resource Control Policies

{% hint style="info" %}
When configuring cross-account access for Bring Your Own Bucket (BYOB), organisational policy layers *Service Control Policies* (SCPs) and *Resource Control Policies* (RCPs) can prevent access even when the bucket policy is valid.
{% endhint %}

For cross-account S3 requests, AWS evaluates permissions in the following order:

1. **Source Account Service Control Policies** determine which actions the principal is allowed to perform, regardless of the destination resource.
2. **Destination Account Resource Control Policies** determine what external principals can do on resources within that account and act as a control layer above the bucket policy.
3. **Bucket Policy and Identity-Based Policies** are evaluated only after both SCP and RCP checks pass. This means that an explicit **Deny**, or the absence of a required **Allow**, in either the SCP or RCP results in an immediate final denial and the bucket policy is not evaluated.

## Configuration

You can use either [IAM User ](/home/h-storage/s-awss3/iam-user-method)or [IAM Role](/home/h-storage/s-awss3/iam-role-method) for setting the permissions with IAM Role offering better security for connecting to your own S3 storage.

#### IAM User

[IAM user](/home/h-storage/s-awss3/iam-user-method) uses **long-term credentials** to connect external systems to your S3 storage. These credentials (access\_key\_id and secret\_access\_key) have to be kept secure and should be regularly rotated, which requires updating the keys in all systems that use these keys.

#### IAM Role

[IAM roles](#iam-role) do not use long-term credentials. Instead temporary (12 hours) security permissions are provided when external systems assume the role. A **permission policy** determines which actions are allowed and a **trust policy** determines who (which software) can assume the role. When Platform Core requests to assume the role, the trust policy is checked to see if Platform Core is allowed to assume the role and if allowed, short-lived credentials are provided so Platform Core can borrow the permissions for that role.

You can enable SSE using an Amazon S3-managed key (SSE-S3). Instructions for using KMS-managed (SSE-KMS) keys are found [here](https://help.ica.illumina.com/home/storage/s-sse-kms.md).

## Considerations

### Synchronization

{% hint style="warning" %}
Because of how [Amazon S3 handles folders](https://docs.aws.amazon.com/AmazonS3/latest/userguide/using-folders.html#delete-folders) and does not send events for S3 folders, the following restrictions must be taken into account for Platform Core project data stored in S3.

* When you create an empty folder in S3, it will not be visible in Platform Core.
* When you move folders in S3, the original, but empty, folder will remain visible in Platform Core and must be manually deleted from there.
* When you delete a folder and its contents in S3, the empty folder will remain visible in Platform Core and must be manually deleted in from there.
* You can not create a project with ./ as prefix since S3 does not allow uploading files with this key prefix.
  {% endhint %}

### S3 region

The AWS S3 bucket must **exist in the same AWS region as the Platform Core project**. See the table below for a mapping of Platform Core project regions to AWS regions:

<table><thead><tr><th width="245">Platform Core Project Region</th><th>AWS Region</th></tr></thead><tbody><tr><td>Australia</td><td>ap-southeast-2</td></tr><tr><td>Canada</td><td>ca-central-1</td></tr><tr><td>Germany</td><td>eu-central-1</td></tr><tr><td>India</td><td>ap-south-1</td></tr><tr><td>Indonesia</td><td>ap-southeast-3</td></tr><tr><td>Israel</td><td>il-central-1</td></tr><tr><td>Japan</td><td>ap-northeast-1</td></tr><tr><td>Singapore</td><td>ap-southeast-1</td></tr><tr><td>South Korea*</td><td>ap-northeast-2</td></tr><tr><td>UK</td><td>eu-west-2</td></tr><tr><td>United Arab Emirates</td><td>me-central-1</td></tr><tr><td>United States</td><td>us-east-1</td></tr></tbody></table>

(\*) BSSH is not currently deployed on the South Korea instance, resulting in limited functionality in this region with regard to sequencer integration.

### Versioned S3 Buckets

You can use **unversioned** (only one copy of an object exists), **versioned** (writing creates new versions) and **suspended** (versioning paused) **buckets** as own S3 storage.

If you connect buckets with object versioning, the data in Platform Core will be automatically synced with the data in object store. When an object is deleted without specifying a particular version, a *Delete marker* is created on the objectstore to indicate that the object has been deleted. Platform Core will reflect the object state by deleting the record from the database. No further action on your side is needed to sync.


# IAM User Method

To use the IAM user method, you need to:

* Set [browser access ](#configure-bucket-cors-permission)to the S3 bucket (CORS).
* Create [data access permissions](#create-data-access-permission-aws-iam-policy) (IAM policy).
* Create the [IAM user](#create-aws-iam-user) and [AWS access key](#create-aws-access-key) and [propagate them](#create-platform-core-storage-credential) to Platform Core.
* Verify S3 [Object Ownership](#s3-object-ownership).
* To use copy and move operations, you need to add the necessary [policy statements](#enabling-cross-account-access-for-copy-and-move-operations) in the S3 bucket policy.

Optional steps:

* It is also best practice to [block public access](#block-public-access-to-s3-bucket-optional) to the S3 bucket.
* If your bucket is KMS-enabled, follow the additional steps described [here](/home/h-storage/s-awss3/s-sse-kms).

{% stepper %}
{% step %}

## Configure Bucket CORS Permission

Platform Core requires **cross-origin resource sharing (CORS) permissions** to write to the S3 bucket for uploads via the browser. Refer to [Configuring cross-origin resource sharing (CORS)](https://docs.aws.amazon.com/AmazonS3/latest/userguide/enabling-cors-examples.html) (expand the *Using the S3 console* section) documentation for instructions on enabling CORS via the **AWS Management Console**.

In the cross-origin resource sharing (CORS) section, enter the following content.

```json
[
    {
        "AllowedHeaders": [
            "*"
        ],
        "AllowedMethods": [
            "HEAD",
            "GET",
            "PUT",
            "POST",
            "DELETE"
        ],
        "AllowedOrigins": [
            "https://ica.illumina.com"
        ],
        "ExposeHeaders": [
            "ETag",
            "x-amz-meta-custom-header"
        ]
    }
]
```

{% endstep %}

{% step %}

## Create Data Access Permission - AWS IAM Policy

Platform Core requires specific permissions to access data in an AWS S3 bucket. These permissions are contained in an **AWS IAM Policy**.

#### Permissions

Refer to the [Creating policies on the JSON tab](https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_create-console.html#access_policies_create-json-editor) documentation for instructions on creating an **AWS IAM Policy via the AWS Management Console**. Use the configuration below during the process, **tab one** shows the code for unversioned buckets, **tab two** the code for versioned and versioning-suspended buckets.

{% tabs %}
{% tab title="Unversioned buckets" %}
Paste the JSON policy document below. Note the example below provides access to all object prefixes in the bucket.

{% hint style="warning" %}
Replace **\<YOUR\_BUCKET\_NAME>** with the name of the S3 bucket you created for Platform Core. Replace **\<YOUR\_FOLDER\_NAME>** with the name of the folder in your S3 bucket.
{% endhint %}

<pre class="language-json"><code class="lang-json"><strong>{
</strong>    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutBucketNotification",
                "s3:ListBucket",
                "s3:GetBucketNotification",
                "s3:GetBucketLocation"
            ],
            "Resource": [
                "arn:aws:s3:::&#x3C;YOUR_BUCKET_NAME>"
            ]
        },
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutObject",
                "s3:GetObject",
                "s3:RestoreObject",
                "s3:DeleteObject",
                "s3:GetObjectTagging",
                "s3:PutObjectTagging"
            ],
            "Resource": "arn:aws:s3:::&#x3C;YOUR_BUCKET_NAME>/&#x3C;YOUR_FOLDER_NAME>/*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sts:GetFederationToken"
            ],
            "Resource": [
                "*"
            ]
        }
    ]
}
</code></pre>

{% endtab %}

{% tab title="Versioned/Suspended Buckets" %}
On **Versioned OR Suspended** buckets, paste the JSON policy document below. Note the example below provides access to all objects prefixes in the bucket.

{% hint style="warning" %}
Replace **\<YOUR\_BUCKET\_NAME>** with the name of the S3 bucket you created for Platform Core. Replace **\<YOUR\_FOLDER\_NAME>** with the name of the folder in your S3 bucket.
{% endhint %}

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutBucketNotification",
                "s3:ListBucket",
                "s3:GetBucketNotification",
                "s3:GetBucketLocation",
                "s3:ListBucketVersions",
                "s3:GetBucketVersioning"
            ],
            "Resource": [
                "arn:aws:s3:::<YOUR_BUCKET_NAME>"
            ]
        },
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutObject",
                "s3:GetObject",
                "s3:RestoreObject",
                "s3:DeleteObject",
                "s3:DeleteObjectVersion",
                "s3:GetObjectVersion",
                "s3:GetObjectTagging",
                "s3:PutObjectTagging",
                "s3:GetObjectVersionTagging",
                "s3:PutObjectVersionTagging"
            ],
            "Resource": "arn:aws:s3:::<YOUR_BUCKET_NAME>/<YOUR_FOLDER_NAME>/*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sts:GetFederationToken"
            ],
            "Resource": [
                "*"
            ]
        }
    ]
}
```

{% endtab %}
{% endtabs %}

#### (Optional) Set policy name to "illumina-core-admin-policy"

To create the **IAM Policy via the AWS CLI,** create a local file named `illumina-core-admin-policy.json` containing the policy content above and run the following command. Be sure the path to the policy document (`--policy-document`) leads to the path where you saved the file:

```bash
aws iam create-policy --policy-name illumina-core-admin-policy --policy-document file://illumina-core-admin-policy.json
```

{% endstep %}

{% step %}

## Create AWS IAM User

An AWS IAM User is needed to create an Access Key for Platform Core to connect to the AWS S3 Bucket. The policy will be attached to the IAM user to grant the user the necessary permissions.

Refer to the [Creating IAM users (console)](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_users_create.html#id_users_create_console) documentation for instructions on **creating an AWS IAM User via the AWS Management Console**. Use the following configuration during the process:

* (optional) Set user name to "illumina\_core\_admin"
* Select the **Programmatic access** option for the type of access.
* Select **Attach existing policies directly** when setting the permissions, and choose the policy created in [Create Data Access Permission - AWS IAM Policy.](#create-data-access-permission-aws-iam-policy)
* (Optional) Retrieve the Access Key ID and Secret Access Key by choosing to **Download .csv.**

To **create the IAM user and attach the policy via the AWS CLI,** enter the following command (AWS IAM users are global resources and do not require a region to be specified). This command creates an IAM user `illumina_core_admin`, retrieves your AWS account number, and then attaches the policy to the user.

```bash
aws iam create-user --user-name illumina_core_admin
ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
aws iam attach-user-policy --policy-arn arn:aws:iam::${ACCOUNT_ID}:policy/illumina-core-admin-policy --user-name illumina_core_admin
```

{% endstep %}

{% step %}

## Create AWS Access Key

{% hint style="warning" %}
If the Access Key information was retrieved during the [IAM user creation](#id-3-create-aws-iam-user), skip this step.
{% endhint %}

Refer to the [Managing access keys (console)](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_credentials_access-keys.html#Using_CreateAccessKey) AWS documentation for instructions on creating an **AWS Access Key via the AWS Console**. See the "To create, modify, or delete another IAM user's access keys (console)" sub-section.

Use the command below to create the Access Key for the illumina\_core\_admin IAM user. Note the `SecretAccessKey` is sensitive and should be stored securely. The access key is only displayed when this command is executed and cannot be recovered. **A new access key must be created if it is lost**.

```bash
aws iam create-access-key --user-name illumina_core_admin

    "AccessKey": {
        "UserName": "illumina_core_admin",
        "AccessKeyId": "<access key id>",
        "Status": "Active",
        "SecretAccessKey": "<secret access key>",
        "CreateDate": "2020-10-22 09:42:24+00:00"
    }
```

The `AccessKeyId` and `SecretAccessKey` values will be provided to Platform Core in the next step.
{% endstep %}

{% step %}

## S3 Bucket Policy

Connecting your S3 bucket to Platform Core does not require any additional bucket policies.

<details>

<summary>What if you need a bucket policy for use cases beyond Platform Core?</summary>

The bucket policy must then support the essential permissions needed by ICA without inadvertently restricting its functionality.

{% hint style="warning" %}
Be sure to replace the following fields:

* YOUR\_BUCKET\_NAME: Replace this field with the name of the S3 bucket you created for Platform Core.
* YOUR\_ACCOUNT\_ID: Replace this field with your account ID number.
* YOUR\_IAM\_USER: Replace this field with the name of your IAM user created for Platform Core.
  {% endhint %}

```json
{
     "Version": "2012-10-17",
     "Statement": [
         {
             "Effect": "Deny",
             "Principal": {
                 "AWS": "*"
             },
             "Action": [
                 "s3:PutObject",
                 "s3:GetObject",
                 "s3:RestoreObject",
                 "s3:DeleteObject",
                 "s3:DeleteObjectVersion",
                 "s3:GetObjectVersion",
                 "s3:GetObjectTagging",
                 "s3:PutObjectTagging",
                 "s3:GetObjectVersionTagging",
                 "s3:PutObjectVersionTagging"
             ],
             "Resource": "arn:aws:s3:::YOUR_BUCKET_NAME/*",
             "Condition": {
                 "ArnNotLike": {
                     "aws:PrincipalArn": [
                         "arn:aws:iam::YOUR_ACCOUNT_ID:user/YOUR_IAM_USER",
                         "arn:aws:sts::YOUR_ACCOUNT_ID:federated-user/*"
                     ]
                 }
             }
         }
     ]
 }
```

In this example, **restriction is enabled** on the bucket policy to prevent any kind of access to the bucket. However, there is an **exception** **rule** added **for the IAM user** that Platform Core is using to connect to the S3 bucket. The exception rule is allowing Platform Core to perform the above S3 action permissions necessary for Platform Core functionalities.

Additionally, the exception rule is applied to the STS federated user session principal associated with Platform Core. Since Platform Core leverages the **AWS STS to provide temporary credentials** that allow users to perform actions on the S3 bucket, it is crucial to include these STS federated user session principals in your policy's whitelist. Failing to do so could result in 403 Forbidden errors when users attempt to interact with the bucket's objects using the provided temporary credentials.

</details>
{% endstep %}

{% step %}

## S3 Object Ownership

Verify that the bucket's ***Object Ownership*****&#x20;is set to&#x20;*****ACLs disabled (Bucket owner enforced)*** on the Permissions tab. This is the AWS-recommended default for new buckets.

If it is not correctly set, open your bucket In the AWS S3 console, go to the **Permissions** tab, find **Object Ownership**, choose **Edit**, select **ACLs disabled (recommended)**, and save.

When Illumina copies data into your bucket, the objects are written by an Illumina service identity. If your bucket has *ACLs enabled (Bucket owner preferred)*, those copied objects remain owned by the writer rather than by your bucket, and Illumina's processing is then unable to read them back, which causes copy jobs to stall or fail with an access-denied (403) error. Setting *Object Ownership* to *ACLs disabled (Bucket owner enforced)* ensures your account always owns every object in the bucket, so the data can be read and processed normally.

{% hint style="info" %}
**Object Ownership is not the same as the Bucket Policy**

*Object Ownership* and *Bucket Policy* are two separate settings on an S3 bucket:

* **Bucket Policy** is the JSON document that grants Illumina access to your bucket.
* **Object Ownership** controls who owns objects written to the bucket.
  {% endhint %}
  {% endstep %}

{% step %}

## Block Public Access to S3 bucket (optional)

By default, public access to the S3 bucket is allowed. For increased security, it is advised to **block public access** with the following command below. Change `<YOUR_BUCKET_NAME>` to the name of your S3 bucket.

```
aws s3api put-public-access-block --bucket <YOUR_BUCKET_NAME> --public-access-block-configuration "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"
```

To block public access to S3 buckets on account level, you can use the AWS Console on the [Amazon Web Services](https://aws.amazon.com/console/) website.
{% endstep %}

{% step %}

## Create Platform Core Storage Credential

### Storage Credential

To connect your S3 account to Platform Core, you need to add a storage credential in Platform Core containing the Access Key ID and Access Key created in the previous step. **From the Platform Core home screen**, navigate to **System Settings > Credentials** > **Create > Storage Credential** to create a new storage credential.

Provide a **name** for the storage credentials, ensure the type is set to "AWS user" and provide the **Access Key ID** and **Secret Access Key**.

{% hint style="info" %}
The key prefix is mandatory in your storage credentials if you created a folder as recommended in step 2 [data access permissions](#id-2-create-data-access-permission-aws-iam-policy).
{% endhint %}

### Storage Configuration

Once the storage credentials are present, create a **storage configuration** using the secret credential. Refer to [Create a Storage Configuration](/home/h-storage#create-a-storage-configuration) for details.
{% endstep %}

{% step %}

## Enabling Cross-Account Access for Copy and Move Operations

{% hint style="warning" %}
When setting up cross-account access, ensure that no [organisational policies](/home/h-storage/s-awss3#service-control-policies-and-resource-control-policies) are blocking the required permissions
{% endhint %}

Platform Core uses **AssumeRole** to copy and move objects from a bucket in an AWS account to another bucket in another AWS account. To allow cross account access to a bucket, the following policy statements must be **added in the S3 bucket policy.**

{% hint style="warning" %}
Be sure to replace the following fields:

* **\<ASSUME\_ROLE\_ARN>**: Replace this field with the ARN of the cross account role you want to give permission to. Refer to the table below to determine which region-specific Role ARN should be used.
* **\<YOUR\_BUCKET\_NAME>**: Replace this field with the name of the S3 bucket you created for Platform Core.
  {% endhint %}

{% tabs %}
{% tab title="Unversioned" %}

```json
  {
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "AllowCrossAccountAccess",
                "Effect": "Allow",
                "Principal": {
                    "AWS": "<ASSUME_ROLE_ARN>"
                },
                "Action": [
                    "s3:PutObject",
                    "s3:DeleteObject",
                    "s3:ListMultipartUploadParts",
                    "s3:AbortMultipartUpload",
                    "s3:GetObject",
                    "s3:GetObjectTagging",
                    "s3:PutObjectTagging"
                ],
                "Resource": [
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>",
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>/*"
                ]
            }
        ]
    }
```

{% endtab %}

{% tab title="Versioned or Suspended" %}

```json
  {
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "AllowCrossAccountAccess",
                "Effect": "Allow",
                "Principal": {
                    "AWS": "<ASSUME_ROLE_ARN>"
                },
                "Action": [
                    "s3:PutObject",
                    "s3:DeleteObject",
                    "s3:ListMultipartUploadParts",
                    "s3:AbortMultipartUpload",
                    "s3:GetObject",
                    "s3:GetObjectVersion",
                    "s3:DeleteObjectVersion",
                    "s3:GetObjectTagging",
                    "s3:PutObjectTagging",
                    "s3:GetObjectVersionTagging",
                    "s3:PutObjectVersionTagging"
                ],
                "Resource": [
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>",
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>/*"
                ]
            }
        ]
    }
```

{% endtab %}
{% endtabs %}

The ARN of the cross account role you want to give permission to is specified in the Principal. Refer to the table below to determine which region-specific Role ARN should be used.

<table><thead><tr><th width="249">Region</th><th>Role ARN</th></tr></thead><tbody><tr><td>Australia (AU)</td><td>arn:aws:iam::079623148045:role/ica_aps2_crossacct</td></tr><tr><td>Canada (CA)</td><td>arn:aws:iam::079623148045:role/ica_cac1_crossacct</td></tr><tr><td>Germany (EU)</td><td>arn:aws:iam::079623148045:role/ica_euc1_crossacct</td></tr><tr><td>India (IN)</td><td>arn:aws:iam::079623148045:role/ica_aps3_crossacct</td></tr><tr><td>Indonesia (ID)</td><td>arn:aws:iam::079623148045:role/ica_aps4_crossacct</td></tr><tr><td>Israel (IL)</td><td>arn:aws:iam::079623148045:role/ica_ilc1_crossacct</td></tr><tr><td>Japan (JP)</td><td>arn:aws:iam::079623148045:role/ica_apn1_crossacct</td></tr><tr><td>Singapore (SG)</td><td>arn:aws:iam::079623148045:role/ica_aps1_crossacct</td></tr><tr><td>South Korea (KR)</td><td>arn:aws:iam::079623148045:role/ica_apn2_crossacct</td></tr><tr><td>Taiwan (TW)</td><td>arn:aws:iam::079623148045:role/ica_ape2_crossacct</td></tr><tr><td>UK (GB)</td><td>arn:aws:iam::079623148045:role/ica_euw2_crossacct</td></tr><tr><td>United Arab Emirates (AE)</td><td>arn:aws:iam::079623148045:role/ica_mec1_crossacct</td></tr><tr><td>United States (US-Oregon)</td><td>arn:aws:iam::079623148045:role/ica_usw2_crossacct</td></tr><tr><td>United States (US-N. Virginia)</td><td>arn:aws:iam::079623148045:role/ica_use1_crossacct</td></tr></tbody></table>

Before copy and move operations are executed **on your own S3 storage**, a test is performed to verify the necessary operational rights. This can result in temporary test files remaining (for example when [IAM policy](#create-data-access-permission-aws-iam-policy) is not correctly set up for a versioned bucket). These files can safely be manually deleted from your S3 console.
{% endstep %}
{% endstepper %}

## Copying Object Tags

{% hint style="info" icon="burst-new" %}
**Storage configurations created in Platform Core 2.47 or later automatically have object tag copying included.**
{% endhint %}

Object tags are copied when performing **Copy**, **Move**, **Archive** and **Unarchive** Operations within the same account or across accounts when using your own S3 storage.

{% hint style="info" %}
For storage configurations created **prior to Platform Core v2.47,** **object tag copying is optional**. If you want to enable it, contact Illumina support to enable TaggingPermissionType on the Platform Core Storage Configuration record associated with the S3 bucket with Object tags.
{% endhint %}

If there is an issue with copying, verify you have the required permission in your policies

* In the configuration above, the **s3:GetObjectTagging** and **s3:PutObjectTagging** are part of the [IAM Policy](#create-data-access-permission-aws-iam-policy).
* **s3:GetObjectTagging** and **s3:PutObjectTagging** are part of the [S3 Bucket policy](#s3-bucket-policy).
* For cross-account copy or move operations, **s3:GetObjectTagging** and **s3:PutObjectTagging** are included in the [cross-account access bucket policy](#enabling-cross-account-access-for-copy-and-move-operations).


# IAM Role Method

To use the IAM Role method, you need to:

* Set [browser access ](#configure-bucket-cors-permission)to the S3 bucket (CORS).
* Create [data access permissions ](#create-data-access-permission-aws-iam-policy)(IAM policy).
* [Configure storage credentials](#create-platform-core-storage-credential) in Platform Core.
* Create the [IAM role](#create-aws-iam-role) and [OIDC provider](#create-openid-connect-oidc-identity-provider).
* [Create a storage configuration](/home/h-storage#create-a-storage-configuration) in Platform Core.
* Verify S3 [Object Ownership](#s3-object-ownership).
* To use copy and move operations, you need to add the necessary policy statements in the S3 bucket policy.

Optionally

* It is best practice to [block public access](#block-public-access-to-s3-bucket-optional) to the S3 bucket.
* If your bucket is KMS-enabled, follow the additional steps described [here](/home/h-storage/s-awss3/s-sse-kms).

{% stepper %}
{% step %}

## Configure Bucket CORS Permission

Platform Core requires **cross-origin resource sharing (CORS) permissions** to write to the S3 bucket for uploads via the browser. Refer to [Configuring cross-origin resource sharing (CORS)](https://docs.aws.amazon.com/AmazonS3/latest/userguide/enabling-cors-examples.html) (expand the *Using the S3 console* section) documentation for instructions on enabling CORS via the **AWS Management Console**.

In the cross-origin resource sharing (CORS) section, enter the following content.

```json
[
    {
        "AllowedHeaders": [
            "*"
        ],
        "AllowedMethods": [
            "HEAD",
            "GET",
            "PUT",
            "POST",
            "DELETE"
        ],
        "AllowedOrigins": [
            "https://ica.illumina.com"
        ],
        "ExposeHeaders": [
            "ETag",
            "x-amz-meta-custom-header"
        ]
    }
]
```

{% endstep %}

{% step %}

## Create Data Access Permission - AWS IAM Policy

Platform Core requires specific permissions to access data in an AWS S3 bucket. These permissions are contained in an **AWS IAM Policy**.

#### Permissions

Refer to the [Creating policies on the JSON tab](https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_create-console.html#access_policies_create-json-editor) documentation for instructions on creating an **AWS IAM Policy via the AWS Management Console (on AWS go to IAM > Policies > create policy)**. Use the configuration below during the process, **tab one** shows the code for unversioned buckets, **tab two** the code for versioned and versioning-suspended buckets.

{% tabs %}
{% tab title="Unversioned buckets" %}
Paste the JSON policy document below. Note the example below provides access to all object prefixes in the bucket.

{% hint style="warning" %}
Replace <**YOUR\_BUCKET\_NAME>** with the name of the S3 bucket you created for Platform Core. Replace <**YOUR\_FOLDER\_NAME>** with the name of the folder in your S3 bucket.
{% endhint %}

<pre class="language-json" data-line-numbers><code class="lang-json"><strong>{
</strong>    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutBucketNotification",
                "s3:ListBucket",
                "s3:GetBucketNotification",
                "s3:GetBucketLocation"
            ],
            "Resource": [
                "arn:aws:s3:::&#x3C;YOUR_BUCKET_NAME>"
            ]
        },
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutObject",
                "s3:GetObject",
                "s3:RestoreObject",
                "s3:DeleteObject",
                "s3:GetObjectTagging",
                "s3:PutObjectTagging"
            ],
            "Resource": "arn:aws:s3:::&#x3C;YOUR_BUCKET_NAME>/&#x3C;YOUR_FOLDER_NAME>/*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sts:GetFederationToken"
            ],
            "Resource": [
                "*"
            ]
        }
    ]
}
</code></pre>

{% endtab %}

{% tab title="Versioned/Suspended Buckets" %}
On **Versioned OR Suspended** buckets, paste the JSON policy document below. Note the example below provides access to all objects prefixes in the bucket.

{% hint style="warning" %}
Replace **YOUR\_BUCKET\_NAME** with the name of the S3 bucket you created for Platform Core. Replace **YOUR\_FOLDER\_NAME** with the name of the folder in your S3 bucket.
{% endhint %}

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutBucketNotification",
                "s3:ListBucket",
                "s3:GetBucketNotification",
                "s3:GetBucketLocation",
                "s3:ListBucketVersions",
                "s3:GetBucketVersioning"
            ],
            "Resource": [
                "arn:aws:s3:::<YOUR_BUCKET_NAME>"
            ]
        },
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutObject",
                "s3:GetObject",
                "s3:RestoreObject",
                "s3:DeleteObject",
                "s3:DeleteObjectVersion",
                "s3:GetObjectVersion",
                "s3:GetObjectTagging",
                "s3:PutObjectTagging",
                "s3:GetObjectVersionTagging",
                "s3:PutObjectVersionTagging"
            ],
            "Resource": "arn:aws:s3:::<YOUR_BUCKET_NAME>/<YOUR_FOLDER_NAME>/*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sts:GetFederationToken"
            ],
            "Resource": [
                "*"
            ]
        }
    ]
}
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
If you get the error "no identity-based policy allows the s3:PutObject action", verify that in the expression above:

* (line 13) "arn:aws:s3:::\<YOUR\_BUCKET\_NAME>" matches the first part of
* (line 26) "arn:aws:s3:::\<YOUR\_BUCKET\_NAME>/\<YOUR\_FOLDER\_NAME>/\*"
  {% endhint %}

#### (Optional) Set policy name to "illumina-core-admin-policy"

To create the **IAM Policy via the AWS CLI,** create a local file named `illumina-core-admin-policy.json` containing the policy content above and run the following command. Be sure the path to the policy document (`--policy-document`) leads to the path where you saved the file:

```bash
aws iam create-policy --policy-name illumina-core-admin-policy --policy-document file://illumina-core-admin-policy.json
```

{% endstep %}

{% step %}

## Create Platform Core Storage Credential

### Storage Credential

To connect your S3 account to Platform Core, you need to add a storage credential in Platform Core which will generate the `RoleSessionName` prefix.

**From the Platform Core home screen**, navigate to **System Settings > Credentials** > **Create > Storage Credential** to create a new storage credential.

1. Select **AWS\_Role** as type and provide a **name** for the storage credential.
2. Choose **Generate** to create the **RoleSessionName** prefix. Once generated, you can **download it** with the Download to Excel button or copy and paste it by unmasking the prefix with the eye symbol on the right. You will need this RoleSessionName in the next step.

{% hint style="warning" %}
**You can only download or copy this value now during creation.** Once this dialog box closes after saving, you can no longer access this value.
{% endhint %}

<figure><img src="/files/OIfLzZo9RVNFJPXdPnx5" alt="" width="563"><figcaption></figcaption></figure>

### Storage Configuration

Once the storage credentials are present, create a **storage configuration** using the credential. Refer to [Create a Storage Configuration](/home/h-storage#create-a-storage-configuration) for details.
{% endstep %}

{% step %}

## Create AWS IAM Role

You need to create the IAM role which Platform Core will assume to access your S3 bucket. See this [AWS documentation](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create.html) for instructions on **creating an AWS IAM Role via the AWS Management Console**. This Role will allow to delegate permissions to Platform Core to connect to your S3 storage for the required duration.

Open your [IAM console](https://console.aws.amazon.com/iam/) and perform the steps below:

1. Copy the **Trust Policy** below to an editor and update the following values:
   * `<your AWS client account number>` is your actual [AWS client account number](https://docs.aws.amazon.com/accounts/latest/reference/manage-acct-identifiers.html#FindingYourAccountIdentifiers).
   * `<region-alias>` must be taken from the [OIDC reference table](#oidc-reference-table) below.\
     For example, us-east-1 for United States.
   * `<OIDC provider ID>` must be taken from the [OIDC reference table](#oidc-reference-table) below.\
     For example 10FBA7EDEB5930CCFC300EF5AC3DB2FE for United States
   * `<namespace>` must be taken from the [OIDC reference table](#oidc-reference-table) below.\
     For example, use1 for United States.
   * `<session name prefix>` must be replaced with the session name prefix value generated in the previous step, [Create Platform Core storage credential](#create-platform-core-storage-credential). **Keep the -\* at the end**.\
     This prefix works as an proof of identity for the requesting process and ensures the role can only be granted if the requesting process provides a session name starting with this prefix. If a process with a different session name prefix requests the role, it will be automatically denied. This is an additional layer of security.
2. Choose **Roles > Create role** and choose **Custom Trust Policy**.
3. Paste the **edited Trust Policy**.
4. Select the **Permission Policy** created in [Create AWS IAM Policy](#create-data-access-permission-aws-iam-policy).
5. Give your role a **name** to indicate what it is to be used for (for example Illumina\_Core\_Role) and preferably a **description** so other users will know what the IAM role will be used for.
6. Click **create** to create the role.
7. Open your created role and choose **Edit** (top right) to set the role **time to 12 hours** instead of the default 1 hour.

{% hint style="warning" %}
If the role time is not set to 12 hours, the storage configuration will not go online.
{% endhint %}

8. **Copy the ARN** from your created role summary as this will be needed in the [Storage Configuration](/home/h-storage#create-a-storage-configuration) in Platform Core

#### Trust Policy

{% hint style="info" %}
Note the double colon symbol (::) before your AWS client account number
{% endhint %}

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "Federated": "arn:aws:iam::<your AWS client account number>:oidc-provider/oidc.eks.<region-alias>.amazonaws.com/id/<OIDC Provider ID>"
            },
            "Action": "sts:AssumeRoleWithWebIdentity",
            "Condition": {
                "StringLike": {
                    "oidc.eks.<region-alias>.amazonaws.com/id/<OIDC Provider ID>:sub": "system:serviceaccount:<namespace>:irsa-<namespace>-gds*",
                    "oidc.eks.<region-alias>.amazonaws.com/id/<OIDC Provider ID>:aud": "sts.amazonaws.com",
                    "sts:RoleSessionName": "<session name prefix>-*"
                }
            }
        }
    ]
    }
```

#### Example

```json
    {
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "Federated": "arn:aws:iam::123456789:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/10FBA7EDEB5930CCFC300EF5AC3DB2FE"
            },
            "Action": "sts:AssumeRoleWithWebIdentity",
            "Condition": {
                "StringLike": {
                    "oidc.eks.us-east-1.amazonaws.com/id/10FBA7EDEB5930CCFC300EF5AC3DB2FE:sub": "system:serviceaccount:use1:irsa-use1-gds*",
                    "oidc.eks.us-east-1.amazonaws.com/id/10FBA7EDEB5930CCFC300EF5AC3DB2FE:aud": "sts.amazonaws.com",
                    "sts:RoleSessionName": "mUlP0AqBmwf9CMpjEUFY7J2z60sdveP-*"
                }
            }
        }
    ]
    }
```

#### OIDC Reference table

{% hint style="info" %}
**Replace \<your AWS client account number> with your actual AWS client account number**
{% endhint %}

<table data-full-width="false"><thead><tr><th width="102.8046875">Region Name</th><th width="137.5078125">OIDC provider ID</th><th width="285.43359375">OIDC Provider ARN</th><th width="116.7421875">Region Alias</th><th width="92.26953125">Namespace</th></tr></thead><tbody><tr><td>Australia (AU)</td><td>F4CD1AEAE6E0820F305F0230FAF6319C</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-southeast-2.amazonaws.com/id/F4CD1AEAE6E0820F305F0230FAF6319C</td><td>ap-southeast-2</td><td>aps2</td></tr><tr><td>Canada (CA)</td><td>4E70F8E1A204A4B2A22E4F7BA9A06D27</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ca-central-1.amazonaws.com/id/4E70F8E1A204A4B2A22E4F7BA9A06D27</td><td>ca-central-1</td><td>cac1</td></tr><tr><td>Germany (EU)</td><td>FD40D1945EBD71D8433A98C0CE04E625</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.eu-central-1.amazonaws.com/id/FD40D1945EBD71D8433A98C0CE04E625</td><td>eu-central-1</td><td>euc1</td></tr><tr><td>India (IN)</td><td>4C3E0D308DB6DA9625FF938C57DAB3B6</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-south-1.amazonaws.com/id/4C3E0D308DB6DA9625FF938C57DAB3B6</td><td>ap-south-1</td><td>aps3</td></tr><tr><td>Indonesia (ID)</td><td>0E0C765DA73BD1FC509FAC71F92BDB5C</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-southeast-3.amazonaws.com/id/0E0C765DA73BD1FC509FAC71F92BDB5C</td><td>ap-southeast-3</td><td>aps4</td></tr><tr><td>Israel (IL)</td><td>EB5CD54864FC17FE53C44E8F9E3943DC</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.il-central-1.amazonaws.com/id/EB5CD54864FC17FE53C44E8F9E3943DC</td><td>il-central-1</td><td>ilc1</td></tr><tr><td>Japan (JP)</td><td>31343141B6F8EA41F379AE795CFA7638</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-northeast-1.amazonaws.com/id/31343141B6F8EA41F379AE795CFA7638</td><td>ap-northeast-1</td><td>apn1</td></tr><tr><td>Singapore (SG)</td><td>4839F25C2D7F1765F0523616EB33711F</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-southeast-1.amazonaws.com/id/4839F25C2D7F1765F0523616EB33711F</td><td>ap-southeast-1</td><td>aps1</td></tr><tr><td>South Korea (KR)</td><td>F2F941225297CB2CD58E91A45ED1362D</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-northeast-2.amazonaws.com/id/F2F941225297CB2CD58E91A45ED1362D</td><td>ap-northeast-2</td><td>apn2</td></tr><tr><td>Taiwan (TW)</td><td>D321ECE5E6F7AEA2F7B0BDE546B9EB39</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.ap-east-2.amazonaws.com/id/D321ECE5E6F7AEA2F7B0BDE546B9EB39</td><td>ap-east-2</td><td>ape2</td></tr><tr><td>United Arab Emirates (UAE)</td><td>183FDEC68B2A1075CBF28D81199C1F3B</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.me-central-1.amazonaws.com/id/183FDEC68B2A1075CBF28D81199C1F3B</td><td>me-central-1</td><td>mec1</td></tr><tr><td>UK (GB)</td><td>CC52F03C88D774F70AB9D2E2BABDF225</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.eu-west-2.amazonaws.com/id/CC52F03C88D774F70AB9D2E2BABDF225</td><td>eu-west-2</td><td>euw2</td></tr><tr><td>United States (US -Oregon)</td><td>8AC895F8C45EF3AE1C7053C56A09C5B9</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.us-west-2.amazonaws.com/id/8AC895F8C45EF3AE1C7053C56A09C5B9</td><td>us-west-2</td><td>usw2</td></tr><tr><td>United States (US - N. Virginia)</td><td>10FBA7EDEB5930CCFC300EF5AC3DB2FE</td><td>arn:aws:iam::&#x3C;your AWS client account number>:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/10FBA7EDEB5930CCFC300EF5AC3DB2FE</td><td>us-east-1</td><td>use1</td></tr></tbody></table>
{% endstep %}

{% step %}

## Create OpenID Connect (OIDC) Identity Provider

An OpenID Connect identity provider is a trusted resource that provides identity tokens. This allows AWS to know **which external identities are allowed to obtain the temporary roles**. Here you connect your regional Platform Core instance so it can obtain the required role to access your storage. For more information on OIDC providers, see [OIDC entity providers](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_providers_create_oidc.html) on AWS.

Open your [IAM console](https://console.aws.amazon.com/iam/) and perform the steps below:

* Under **IAM > Identity Providers > Add provider**.
* Select **OpenID Connect** as provider and enter the **Provider URL** which matches your Platform Core/S3 location from the [table](#oidc-provider-locations) below.
* For **Audience**, enter **sts.amazonaws.com** and click on **Add Provider.**
* **Verify** that the **arn** from the newly created OIDC provider matches the arn from the [Trust Policy](#trust-policy) above.

See below for an example of how the OIDC provider and IAM role Trusted entities look

<figure><img src="/files/Rv2VTc3erbtZar6ubTk0" alt=""><figcaption></figcaption></figure>

#### OIDC Provider Locations

<table><thead><tr><th width="171.8046875">Region</th><th>OIDC Provider URL</th></tr></thead><tbody><tr><td>Australia (AU)</td><td>https://oidc.eks.ap-southeast-2.amazonaws.com/id/F4CD1AEAE6E0820F305F0230FAF6319C</td></tr><tr><td>Canada (CA)</td><td>https://oidc.eks.ca-central-1.amazonaws.com/id/4E70F8E1A204A4B2A22E4F7BA9A06D27</td></tr><tr><td>Germany (EU)</td><td>https://oidc.eks.eu-central-1.amazonaws.com/id/FD40D1945EBD71D8433A98C0CE04E625</td></tr><tr><td>India (IN)</td><td>https://oidc.eks.ap-south-1.amazonaws.com/id/4C3E0D308DB6DA9625FF938C57DAB3B6</td></tr><tr><td>Indonesia (ID)</td><td>https://oidc.eks.ap-southeast-3.amazonaws.com/id/0E0C765DA73BD1FC509FAC71F92BDB5C</td></tr><tr><td>Israel (IL)</td><td>https://oidc.eks.il-central-1.amazonaws.com/id/EB5CD54864FC17FE53C44E8F9E3943DC</td></tr><tr><td>Japan (JP)</td><td>https://oidc.eks.ap-northeast-1.amazonaws.com/id/31343141B6F8EA41F379AE795CFA7638</td></tr><tr><td>Singapore (SG)</td><td>https://oidc.eks.ap-southeast-1.amazonaws.com/id/4839F25C2D7F1765F0523616EB33711F</td></tr><tr><td>South Korea (KR)</td><td>https://oidc.eks.ap-northeast-2.amazonaws.com/id/F2F941225297CB2CD58E91A45ED1362D</td></tr><tr><td>Taiwan (TW)</td><td>https://oidc.eks.ap-east-2.amazonaws.com/id/D321ECE5E6F7AEA2F7B0BDE546B9EB39</td></tr><tr><td>United Arab Emirates (UAE)</td><td>https://oidc.eks.me-central-1.amazonaws.com/id/183FDEC68B2A1075CBF28D81199C1F3B</td></tr><tr><td>UK (GB)</td><td>https://oidc.eks.eu-west-2.amazonaws.com/id/CC52F03C88D774F70AB9D2E2BABDF225</td></tr><tr><td>United States (US - Oregon)</td><td>https://oidc.eks.us-west-2.amazonaws.com/id/8AC895F8C45EF3AE1C7053C56A09C5B9</td></tr><tr><td>United States (US - N. Virginia)</td><td>https://oidc.eks.us-east-1.amazonaws.com/id/10FBA7EDEB5930CCFC300EF5AC3DB2FE</td></tr></tbody></table>
{% endstep %}

{% step %}

## S3 Bucket Policy

Connecting your S3 bucket to Platform Core does not require any additional bucket policies.

<details>

<summary>What if you need a bucket policy for use cases beyond Platform Core?</summary>

The bucket policy must then support the essential permissions needed by Platform Core without inadvertently restricting its functionality.

{% hint style="warning" %}
Be sure to replace the following fields:

* YOUR\_BUCKET\_NAME: Replace this field with the name of the S3 bucket you created for Platform Core.
* YOUR\_ACCOUNT\_ID: Replace this field with your account ID number.
* YOUR\_IAM\_ROLE: Replace this field with the name of your IAM role created for Platform Core.
  {% endhint %}

```json
{
     "Version": "2012-10-17",
     "Statement": [
         {
             "Effect": "Deny",
             "Principal": {
                 "AWS": "*"
             },
             "Action": [
                 "s3:PutObject",
                 "s3:GetObject",
                 "s3:RestoreObject",
                 "s3:DeleteObject",
                 "s3:DeleteObjectVersion",
                 "s3:GetObjectVersion",
                 "s3:GetObjectTagging",
                 "s3:PutObjectTagging",
                 "s3:GetObjectVersionTagging",
                 "s3:PutObjectVersionTagging"
             ],
             "Resource": "arn:aws:s3:::YOUR_BUCKET_NAME/*",
             "Condition": {
                 "ArnNotLike": {
                     "aws:PrincipalArn": [
                         "arn:aws:iam::YOUR_ACCOUNT_ID:role/YOUR_IAM_ROLE",
                         "arn:aws:sts::YOUR_ACCOUNT_ID:federated-user/*"
                     ]
                 }
             }
         }
     ]
 }
```

In this example, **restriction is enabled** on the bucket policy to prevent any kind of access to the bucket. However, there is an **exception** **rule** added **for the IAM role** that Platform Core is using to connect to the S3 bucket. The exception rule is allowing Platform Core to perform the above S3 action permissions necessary for Platform Core functionalities.

Additionally, the exception rule is applied to the STS federated user session principal associated with Platform Core. Since Platform Core leverages the **AWS STS to provide temporary credentials** that allow users to perform actions on the S3 bucket, it is crucial to include these STS federated user session principals in your policy's whitelist. Failing to do so could result in 403 Forbidden errors when users attempt to interact with the bucket's objects using the provided temporary credentials.

</details>
{% endstep %}

{% step %}

## S3 Object Ownership

Verify that the bucket's ***Object Ownership*****&#x20;is set to&#x20;*****ACLs disabled (Bucket owner enforced)*** on the Permissions tab. This is the AWS-recommended default for new buckets.

If it is not correctly set, open your bucket In the AWS S3 console, go to the **Permissions** tab, find **Object Ownership**, choose **Edit**, select **ACLs disabled (recommended)**, and save.

When Illumina copies data into your bucket, the objects are written by an Illumina service identity. If your bucket has *ACLs enabled (Bucket owner preferred)*, those copied objects remain owned by the writer rather than by your bucket, and Illumina's processing is then unable to read them back, which causes copy jobs to stall or fail with an access-denied (403) error. Setting *Object Ownership* to *ACLs disabled (Bucket owner enforced)* ensures your account always owns every object in the bucket, so the data can be read and processed normally.

{% hint style="info" %}
**Object Ownership is not the same as the Bucket Policy**

*Object Ownership* and *Bucket Policy* are two separate settings on an S3 bucket:

* **Bucket Policy** is the JSON document that grants Illumina access to your bucket.
* **Object Ownership** controls who owns objects written to the bucket.
  {% endhint %}
  {% endstep %}

{% step %}

## Block Public Access to S3 bucket (optional)

By default, public access to the S3 bucket is allowed. For increased security, it is advised to **block public access** with the command below. Change `<YOUR_BUCKET_NAME>` to the name of your S3 bucket.

```
aws s3api put-public-access-block --bucket <YOUR_BUCKET_NAME> --public-access-block-configuration "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"
```

To block public access to S3 buckets on account level, you can use the AWS Console on the [Amazon Web Services](https://aws.amazon.com/console/) website.
{% endstep %}

{% step %}

## Enabling Cross-Account Access for Copy and Move Operations

{% hint style="warning" %}
When setting up cross-account access, ensure that no [organisational policies](/home/h-storage/s-awss3#service-control-policies-and-resource-control-policies) are blocking the required permissions
{% endhint %}

Platform Core uses **AssumeRole** to copy and move objects from a bucket in an AWS account to another bucket in another AWS account. To allow cross account access to a bucket, the following policy statements must be **added in the S3 bucket policy (tab one** below shows the code for unversioned buckets, **tab two** the code for versioned and versioning-suspended buckets.)

{% hint style="warning" %}
Be sure to replace the following fields:

* **\<ASSUME\_ROLE\_ARN>**: Replace this field with the ARN of the cross account role you want to give permission to. Refer to the table below to determine which region-specific Role ARN should be used.
* **\<YOUR\_BUCKET\_NAME>**: Replace this field with the name of the S3 bucket you created for Platform Core.
  {% endhint %}

{% tabs %}
{% tab title="Unversioned" %}

```json
  {
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "AllowCrossAccountAccess",
                "Effect": "Allow",
                "Principal": {
                    "AWS": "<ASSUME_ROLE_ARN>"
                },
                "Action": [
                    "s3:PutObject",
                    "s3:DeleteObject",
                    "s3:ListMultipartUploadParts",
                    "s3:AbortMultipartUpload",
                    "s3:GetObject",
                    "s3:GetObjectTagging",
                    "s3:PutObjectTagging"
                ],
                "Resource": [
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>",
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>/*"
                ]
            }
        ]
    }
```

{% endtab %}

{% tab title="Versioned or Suspended" %}

```json
  {
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "AllowCrossAccountAccess",
                "Effect": "Allow",
                "Principal": {
                    "AWS": "<ASSUME_ROLE_ARN>"
                },
                "Action": [
                    "s3:PutObject",
                    "s3:DeleteObject",
                    "s3:ListMultipartUploadParts",
                    "s3:AbortMultipartUpload",
                    "s3:GetObject",
                    "s3:GetObjectVersion",
                    "s3:DeleteObjectVersion",
                    "s3:GetObjectTagging",
                    "s3:PutObjectTagging",
                    "s3:GetObjectVersionTagging",
                    "s3:PutObjectVersionTagging"
                ],
                "Resource": [
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>",
                    "arn:aws:s3:::<YOUR_BUCKET_NAME>/*"
                ]
            }
        ]
    }
```

{% endtab %}
{% endtabs %}

The ARN of the cross account role you want to give permission to is specified in the Principal. Refer to the table below to determine which region-specific Role ARN should be used.

<table><thead><tr><th width="249.5252685546875">Region</th><th>Role ARN</th></tr></thead><tbody><tr><td>Australia (AU)</td><td>arn:aws:iam::079623148045:role/ica_aps2_crossacct</td></tr><tr><td>Canada (CA)</td><td>arn:aws:iam::079623148045:role/ica_cac1_crossacct</td></tr><tr><td>Germany (EU)</td><td>arn:aws:iam::079623148045:role/ica_euc1_crossacct</td></tr><tr><td>India (IN)</td><td>arn:aws:iam::079623148045:role/ica_aps3_crossacct</td></tr><tr><td>Indonesia (ID)</td><td>arn:aws:iam::079623148045:role/ica_aps4_crossacct</td></tr><tr><td>Israel (IL)</td><td>arn:aws:iam::079623148045:role/ica_ilc1_crossacct</td></tr><tr><td>Japan (JP)</td><td>arn:aws:iam::079623148045:role/ica_apn1_crossacct</td></tr><tr><td>Singapore (SG)</td><td>arn:aws:iam::079623148045:role/ica_aps1_crossacct</td></tr><tr><td>South Korea (KR)</td><td>arn:aws:iam::079623148045:role/ica_apn2_crossacct</td></tr><tr><td>Taiwan (TW)</td><td>arn:aws:iam::079623148045:role/ica_ape2_crossacct</td></tr><tr><td>United Arab Emirates (AE)</td><td>arn:aws:iam::079623148045:role/ica_mec1_crossacct</td></tr><tr><td>UK (GB)</td><td>arn:aws:iam::079623148045:role/ica_euw2_crossacct</td></tr><tr><td>United States (US - Oregon)</td><td>arn:aws:iam::079623148045:role/ica_usw2_crossacct</td></tr><tr><td>United States (US - N. Virginia)</td><td>arn:aws:iam::079623148045:role/ica_use1_crossacct</td></tr></tbody></table>

Before copy and move operations are executed **on your own S3 storage**, a test is performed to verify the necessary operational rights. This can result in temporary test files remaining (for example when [IAM policy](#create-data-access-permission-aws-iam-policy) is not correctly set up for a versioned bucket). These files can safely be manually deleted from your S3 console.
{% endstep %}
{% endstepper %}

## Copying Object Tags

{% hint style="info" icon="burst-new" %}
**Storage configurations created in Platform Core 2.47 or later automatically have object tag copying included.**
{% endhint %}

Object tags are copied when performing **Copy**, **Move**, **Archive** and **Unarchive** Operations within the same account or across accounts when using your own S3 storage.

{% hint style="info" %}
For storage configurations created **prior to Platform Core v2.47,** **object tag copying is optional**. If you want to enable it, contact Illumina support to enable TaggingPermissionType on the Platform Core Storage Configuration record associated with the S3 bucket with Object tags.
{% endhint %}

If there is an issue with copying, verify you have the required permission in your policies

* In the configuration above, the **s3:GetObjectTagging** and **s3:PutObjectTagging** are part of the [IAM Policy](#create-data-access-permission-aws-iam-policy).
* **s3:GetObjectTagging** and **s3:PutObjectTagging** are part of the [S3 Bucket policy](#s3-bucket-policy).
* For cross-account copy or move operations, **s3:GetObjectTagging** and **s3:PutObjectTagging** are included in the [cross-account access bucket policy](#enabling-cross-account-access-for-copy-and-move-operations).

## Troubleshooting

The table below show some typical error situations and how to resolve them. After performing the configuration update suggested below, perform the validate action (**System Settings > Storage > select your storage > Manage > Validate**) to quickly see if this has resolved your issue.

<table><thead><tr><th width="258.82421875">Error</th><th>Possible Solution</th></tr></thead><tbody><tr><td>GetTempraryCredentaislAsync Failed with STS error</td><td>Not Authorized to perform sts:AssumeRoleWithWebIdentity can indicate an error in the address of the service account that issued the token. Please verify the line <code>"oidc.eks.&#x3C;region-alias>.amazonaws.com/id/&#x3C;OIDC Provider ID>:sub": "system:serviceaccount:&#x3C;namespace>:irsa-&#x3C;namespace>-gds*",</code> for errors in the namespace of your <a href="#create-aws-iam-role">IAM Role</a>.</td></tr><tr><td>Invalid Role Session Duration Set Maximum session Duration to 12 hours.</td><td>This error indicates an incorrect <a href="#create-aws-iam-role">IAM Role</a> session duration. By default it is 1 hour, but It must be set to 12 hours.</td></tr><tr><td>Access Forbidden not authorized to perform s3:GetBucketLocation</td><td>If the cause is no identity-based policy allows the s3:GetBucketLocation action, the issue might be that the permission policy attached to your role does not point to the correct bucket. Please verify the line <code>"Resource": ["arn:aws:s3:::YOUR_BUCKET_NAME"]</code> has the correct bucket name in the <a href="#create-data-access-permission-aws-iam-policy">IAM Policy</a></td></tr></tbody></table>


# SSE-KMS Encryption

This section describes how to connect an AWS S3 Bucket with [SSE-KMS Encryption](https://docs.aws.amazon.com/AmazonS3/latest/userguide/UsingKMSEncryption.html) enabled. General instructions for configuring your AWS account to allow Platform Core to connect to an S3 bucket are found on [this page](/home/h-storage/s-awss3).

{% embed url="<https://www.youtube.com/watch?v=CrcZ5GtSMSY>" %}
Connect an AWS S3 Bucket with SSE-KMS Encryption Enabled
{% endembed %}

## Create an S3 bucket with SSE-KMS

Follow the [AWS instructions](https://docs.aws.amazon.com/AmazonS3/latest/userguide/configuring-bucket-key.html) for how to create S3 bucket with SSE-KMS key.

{% hint style="warning" %}
S3-SSE-KMS must be in the same region as your Platform Core project. See the [Platform Core S3 bucket documentation ](/home/h-storage/s-awss3)for more information.
{% endhint %}

In the "Default encryption" section, enable Server-side encryption and choose `AWS Key Management Service key (SSE-KMS)`. Then select `Choose your AWS KMS key`.

{% hint style="info" %}
If you do not have an existing customer managed key, click `Create a KMS key` and follow [these steps](https://docs.aws.amazon.com/kms/latest/developerguide/create-keys.html) from AWS.
{% endhint %}

<figure><img src="/files/NSHSV2zjJG47z1fgHdJx" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
Once the bucket is set, create a folder with encryption enabled in the bucket that will be linked in the Platform Core storage configuration. This folder will be connected to Platform Core as a [prefix](#create-the-s3-sse-kms-configuration-in-platform-core). Although it is technically possible to use the **root folder**, this **is not recommended** as it will cause the S3 bucket to no longer be available for other projects.
{% endhint %}

![sse-kms-1](/files/1P8CVXEh98PvVyhOOENK)

## Connect the S3-SSE-KMS to Platform Core

Follow the [general instructions ](/home/h-storage/s-awss3)for connecting an S3 bucket to Platform Core.

In the step [Create AWS IAM Policy (IAM User)](/home/h-storage/s-awss3/iam-user-method#create-data-access-permission-aws-iam-policy) or [Create AWS IAM Policy (IAM Role)](/home/h-storage/s-awss3/iam-role-method#create-data-access-permission-aws-iam-policy) update the following:

* Add permission to use KMS key by adding `kms:Decrypt`, `kms:Encrypt`, and `kms:GenerateDataKey`
* Add the ARN KMS key `arn:aws:kms:xxx` on the first "Resource"
* Depending on the bucket type (Unversioned, Versioned or Suspended) the permissions must match the following.

{% tabs %}
{% tab title="Unversioned" %}

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "kms:Decrypt",
                "kms:Encrypt",
                "kms:GenerateDataKey",
                "s3:PutBucketNotification",
                "s3:ListBucket",
                "s3:GetBucketNotification",
                "s3:GetBucketLocation"
            ],
            "Resource": [
                "arn:aws:kms:xxx",
                "arn:aws:s3:::YOUR_BUCKET_NAME"
            ]
        },
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutObject",
                "s3:GetObject",
                "s3:RestoreObject",
                "s3:DeleteObject",
                "s3:GetObjectTagging",
                "s3:PutObjectTagging"
            ],
            "Resource": "arn:aws:s3:::YOUR_BUCKET_NAME/YOUR_FOLDER_NAME/*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sts:GetFederationToken"
            ],
            "Resource": [
                "*"
            ]
        }
    ]
}
```

{% endtab %}

{% tab title="Versioned or Suspended" %}

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "kms:Decrypt",
                "kms:Encrypt",
                "kms:GenerateDataKey",
                "s3:PutBucketNotification",
                "s3:ListBucket",
                "s3:GetBucketNotification",
                "s3:GetBucketLocation",
                "s3:ListBucketVersions",
                "s3:GetBucketVersioning"
            ],
            "Resource": [
                "arn:aws:kms:xxx",
                "arn:aws:s3:::YOUR_BUCKET_NAME"
            ]
        },
        {
            "Effect": "Allow",
            "Action": [
                "s3:PutObject",
                "s3:GetObject",
                "s3:RestoreObject",
                "s3:DeleteObject",
                "s3:DeleteObjectVersion",
                "s3:GetObjectVersion",
                "s3:GetObjectTagging",
                "s3:PutObjectTagging",
                "s3:GetObjectVersionTagging",
                "s3:PutObjectVersionTagging"
            ],
            "Resource": "arn:aws:s3:::YOUR_BUCKET_NAME/YOUR_FOLDER_NAME/*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sts:GetFederationToken"
            ],
            "Resource": [
                "*"
            ]
        }
    ]
}
```

{% endtab %}
{% endtabs %}

At the end of the policy setting, there should be 3 permissions listed in the "Summary".

![sse-kms-2](/files/DrCRzpxhHa8Dg1fAsfKl)

## Create S3-SSE-KMS configuration in Platform Core

Follow the [general instructions](/home/h-storage#create-a-storage-configuration) for how to create a storage configuration in Platform Core.

On step 3 in process above, continue with the `[Optional] Server Side Encryption` to enter the algorithm and key name for server-side encryption processes.

* On "Algorithm", input `aws:kms`
* On "Key Name", input the ARN KMS key: `arn:aws:kms:xxx`

{% hint style="warning" %}
Although "Key prefix" is optional, it is highly recommended to use this and not use the root folder of your S3 bucket. "Key prefix" refers to the folder name in the bucket which you created.

Once a key prefix is used in a storage configuration, no additional storage configurations can be created under that same path.
{% endhint %}

<figure><img src="/files/WtuonlfyA0MEkNXjVFZN" alt="" width="563"><figcaption></figcaption></figure>

## Cross-Account Copy Setup for S3 buckets with SSE-KMS encryption

### KMS Policy

In addition to following the instructions to [Enable Cross-Account Access (IAM User)](/home/h-storage/s-awss3/iam-user-method#enabling-cross-account-access-for-copy-and-move-operations) and [Enable Cross-Account Access (IAM Role)](/home/h-storage/s-awss3/iam-role-method#enabling-cross-account-access-for-copy-and-move-operations), the **KMS policy** must include the following statement for AWS S3 Bucket with SSE-KMS Encryption (refer to the Role ARN table from the IAM user or IAM role page for the `ASSUME_ROLE_ARN` value):

```json
    {
        "Sid": "AllowCrossAccountAccess",
        "Effect": "Allow",
        "Principal": {
            "AWS": "ASSUME_ROLE_ARN"
        },
        "Action": [
            "kms:Encrypt",
            "kms:Decrypt",
            "kms:ReEncrypt*",
            "kms:GenerateDataKey*",
            "kms:DescribeKey"
        ],
        "Resource": "*"
    }
```

## Key Update

If your bucket uses *SSE-KMS* encryption with a self-managed key and you want to update your key, two things must stay in sync for Illumina to read your data:

* **The key must match**: the key configured in your Illumina volume / storage configuration must be the same key that is actually used to encrypt the objects in your bucket.
* **Decrypt permission must be retained**: the identity Illumina uses to read your objects must keep permission to decrypt with that key (the key policy must grant kms:Decrypt to that account/role).

Changing your bucket's default encryption key, rotating to a new key, or restricting the key policy *without updating your Illumina configuration* can cause **access-denied errors** on files that were otherwise copied correctly.


# Troubleshooting AWS-S3 Connectivity

## Common Issues

The following are common issues encountered when connecting an AWS S3 bucket through a storage configuration

<table><thead><tr><th width="149.01739501953125">Error Type</th><th>Error Message</th><th>Description/Fix</th></tr></thead><tbody><tr><td>Access Forbidden</td><td>Access forbidden: {message}</td><td>Mostly occurs because of lack of permission. Fix: Review IAM policy, Bucket policy, ACLs for required permissions</td></tr><tr><td>Unsupported principal</td><td>Unsupported principal: The policy type ${policy_type} does not support the Principal element. Remove the Principal element.</td><td>This can indicate that the <a href="/pages/GqOaepZs9W6aWWyZEOIp#id-6-s3-bucket-policy">S3 bucket policy</a> settings have been added to the <a href="/pages/HpAJcCvQKMhyOjYoNTJo#id-2-create-data-access-permission-aws-iam-policy">IAM policy</a> by mistake.</td></tr><tr><td>Conflict</td><td>System topic is not in a valid state</td><td></td></tr><tr><td>Conflict</td><td>Found conflicting storage container notifications with overlapping prefixes</td><td>See <a href="#conflicting-bucket-notifications">Conflicting bucket notifications</a></td></tr><tr><td>Conflict</td><td>Found conflicting storage container notifications for {prefix}{eventTypeMsg}</td><td>See <a href="#conflicting-bucket-notifications">Conflicting bucket notifications</a></td></tr><tr><td>Conflict</td><td>Found conflicting storage container notifications with overlapping prefixes{prefixMsg}{eventTypeMsg}</td><td>See <a href="#conflicting-bucket-notifications">Conflicting bucket notifications</a></td></tr><tr><td>Customer Container Notification Exists</td><td>Volume Configuration cannot be provisioned: storage container is already set up for customer's own notification</td><td>See <a href="#conflicting-bucket-notifications">Conflicting bucket notifications</a></td></tr><tr><td>Invalid Access Key ID</td><td>Failed to update bucket policy: The AWS Access Key Id you provided does not exist in our records.</td><td>Check the status of the AWS Access Key ID in the console. If not active, activate it. If missing, create it.</td></tr><tr><td>Invalid Paramater</td><td>Missing credentials for storage container</td><td>Check the storage credential. AccessKeyId and/or SecretAccessKey is not set.</td></tr><tr><td>Invalid Parameter</td><td>Missing bucket name for storage container</td><td>Bucket name has not been set for the storage configuration.</td></tr><tr><td>Invalid Parameter</td><td>The storage container name has invalid characters</td><td>Storage container name can only contain lowercase letters, numbers, hyphens, and periods.</td></tr><tr><td>Invalid Parameter</td><td>Storage Container '{storageContainer}' does not exist</td><td>Update storage configuration container to a valid s3 bucket.</td></tr><tr><td>Invalid Parameter</td><td>Invalid parameters for volume configuration: {message}</td><td></td></tr><tr><td>Invalid Storage Container Location</td><td>Storage container must be located in the {region} region</td><td>Update storage configuration region to match storage container region.</td></tr><tr><td>Invalid Storage Container Location</td><td>Storage container must be located in one of the following regions: {regions}</td><td>Update storage configuration region to match storage container region.</td></tr><tr><td>Incorrect bucket <em>Object Ownership</em></td><td>A <strong>copy job</strong> to your bucket does not complete (appears <strong>stuck</strong>), or <strong>fails</strong> with a 403 / access-denied error, even though the connection/health check for your bucket passes.</td><td><ol><li>In the AWS S3 console, set the bucket's <em><strong>Object Ownership</strong></em> to <em><strong>ACLs disabled</strong> (Bucket owner enforced)</em> via the Permissions tab, then Object Ownership, then Edit.</li><li>Re-run the copy job.</li></ol><p><em>Note:</em> Changing this setting fixes all <em>new</em> copies. <strong>Files that were already copied while the setting was incorrect are not retroactively fixed and must be copied again</strong>.</p></td></tr><tr><td>Missing Configuration</td><td>Missing queue name for storage container notification</td><td></td></tr><tr><td>Missing Configuration</td><td>Missing system topic name for storage container notification</td><td></td></tr><tr><td>Missing Configuration</td><td>Missing lambda ARN for storage container notification</td><td></td></tr><tr><td>Missing Configuration</td><td>Missing subscription name for storage container notification</td><td></td></tr><tr><td>Missing Storage Account Settings</td><td>The storage account '{storageAccountName}' needs HNS (Hierarchical Namespace) enabled.</td><td></td></tr><tr><td>Missing Storage Container Settings</td><td>Missing settings for storage container</td><td></td></tr></tbody></table>

## Specific Errors

### Conflicting bucket notifications

This error occurs when an existing bucket notification's event information overlaps with the notifications Platform Core is trying to add. [Amazon S3 event notification](https://docs.aws.amazon.com/AmazonS3/latest/userguide/notification-how-to-event-types-and-destinations.html) only allows overlapping events with non-overlapping prefix. Depending on the conflicts on the notifications, the error can be presented in any of the following:

* *Volume Configuration cannot be provisioned: storage container is already set up for customer's own notification.*
* *Invalid parameters for volume configuration: found conflicting storage container notifications with overlapping prefixes.*
* *Failed to update bucket policy: Configurations overlap. Configurations on the same bucket cannot share a common event type.*

*Solution:*

1. In the Amazon S3 Console, review your current S3 bucket's notification configuration and look for prefixes that overlap with your Storage Configuration's key prefix.
2. Delete the existing notification that overlaps with your Storage Configuration's key prefix.
3. Platform Core will perform a series of steps in the background to re-verify the connection to your bucket.

### GetTemporaryUploadCredentialsAsync failure

This error can occur when recreating a recently deleted storage configuration.\
To fix the issue, you have to delete the bucket notifications:

1. In the [Amazon S3 Console](https://console.aws.amazon.com/s3/) **select the bucket** for which you need to delete the notifications from the list.
2. Choose **properties**.
3. Navigate to the **Event Notifications** section and choose the check box for the event notifications with name *gds:objectcreated*, *gds:objectremoved* and *gds:objectrestore* and click Delete.
4. revalidate the current storage configuration for an immediate update on the **System Settings > Storage > Manage > Validate.**

{% hint style="info" %}
If you do not want to wait revalidate, you can wait 15 minutes, for the storage to become available in Platform Core.
{% endhint %}


# Data

The Data inventory provides access to the files and folders stored in the project or linked to the project. Here, you can perform searches and data management operations such as moving, copying, deleting and (un)archiving.

{% hint style="info" %}
See also [Non-indexed folders ](/project/p-data/non-indexed-folders)which are a special form of data storage optimised for fast processing.
{% endhint %}

### Recommended Practices

#### File/Folder Naming

Platform Core supports UTF-8 characters in file and folder names for data. Please follow the guidelines detailed below. (For more information about recommended approaches to file naming that can be applicable across platforms, please refer to the [AWS S3 documentation](https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-keys.html).)

{% hint style="info" %}
**Folders and files cannot be renamed** after they have been created. **To rename a folder**, you will need to create a new folder with the desired name, move the contents from the original folder into the new one, and then delete the original folder. Please see the [Moving Data](#moving-data) section for more information.
{% endhint %}

<details>

<summary>Characters generally considered "safe"</summary>

* Alphanumeric characters
  * 0-9
  * a-z
  * A-Z
* Special characters
  * Exclamation point `!`
  * Hyphen `-`
  * Underscore `_`
  * Period `.`
  * Asterisk `*`
  * Single quote `'`
  * Open parenthesis `(`
  * Closed parenthesis `)`

</details>

#### Data Formats

See the list of supported [Data Formats](/reference/r-dataformats)

#### Data Privacy

When adding data to Platform Core, prioritize data privacy. Whether you're using storage configurations like AWS S3 or performing Platform Core data uploads, careful management of data access needs to be considered. When setting up cloud storage, confirm that configuration settings prevent unauthorized access. Always verify that uploads are free of unintended data to avoid privacy breaches. For more detailed information, refer to the Platform Core [Security and Compliance ](/reference/r-securityandcompliance)section.

#### Data Integrity

See [Data Integrity](/project/p-data/data-integrity)

***

## Viewing Data

On the **Projects > your\_project > Data** page, you can view file information and preview files.

### Folder view (structured) vs flat view (list)

You can switch between **folder** view and **flat** view with the icons at the left.

* **Folder view** shows the navigation structure and only the **files and folders in the** **current folder**.\
  **Searches in folder view** will be performed on **the** **current folder and all subfolders of the current folder.**
* **Flat view** shows a list of **all files and folders** within the current project.\
  When you perform **searches** in flat view, **all data of your project** will be considered.

**Tree view** shows the navigation structure and only the **files and folders in the** **current folder**.\
**Searches in tree view** will be performed on **the** **current folder and all subfolders of the current folder.**

<figure><img src="/files/l9MqElgSyLk9BzWwgOX5" alt=""><figcaption></figcaption></figure>

**List view** shows **all files and folders** within the current project.\
When you perform **searches** in list view, **all data of your project** will be considered.

<figure><img src="/files/eSLDYOeM4EWlXu4VCDUt" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
You **cannot switch between folder and flat view when you are viewing search results**. Clear the search with the clear search button, or the x in the search dialog first.
{% endhint %}

### Files

To view **file details** click on the filename to see the file details.

* **Run input tags** identifies the last 100 pipelines which used this file as input.
* **Run output tags** identifies the pipeline which created the file.
* **Connector tags** show if the file was added via browser upload or connector.
* Clicking on a **folder** will open the folder itself, to see the folder details, use the **folder details** link at the top right of the screen.

To view **file contents,** select the checkbox at the beginning of the line and then select **View** from the top menu. Alternatively, you can first click on the filename to see the details and then click the view tab to preview the file.

If your data is the **result of an analysis**, you can find the analysis which created it at **Projects > your\_project > Data > your\_data > view > Data details tab > Source analysis**. Clicking the link here will open the analysis.

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th></th><th></th></tr></thead><tbody><tr><td><strong>Filtering</strong></td><td>To add filters, select the <strong>funnel</strong>/<strong>filter</strong> <strong>symbol</strong> at the top right, next to the search field.</td><td>Filters are <strong>reset</strong> when you <strong>exit the current screen</strong>.</td><td></td><td></td></tr><tr><td><strong>Sorting</strong></td><td>To sort data, select the <strong>three vertical dots</strong> in the column header on which you want to sort and chose ascending or descending.</td><td>Sorting is <strong>retained</strong> when you <strong>exit the current screen</strong>.</td><td></td><td></td></tr><tr><td><strong>Displaying Columns</strong></td><td>To change which columns are displayed, select the <strong>three</strong> <strong>columns</strong> <strong>symbol</strong> and select which columns should be shown.</td><td>In <a href="/pages/tRbg9eiA1JFbuoZeBVoH#externally-managed-projects">externally-managed projects</a>, you can see which files are externally controlled and which are ICA-managed by means of the “<strong>managed by</strong>” column.</td><td>In <a href="#tree-view-vs-list-view">flat</a> view, you can jump to the folder in which files are located with the "<strong>Path</strong>" column.</td><td>The displayed columns are <strong>retained</strong> when you <strong>exit the current</strong> screen.</td></tr></tbody></table>

{% hint style="info" %}
When you share the data view by sharing the link from your browser, filters and sorting is retained in links, so the recipient will see the same data and order.
{% endhint %}

To see the **ongoing actions** (copying and moving) on data in the data overview (**Projects > your\_project > Data**), add the **ongoing actions** column from the column list if it is not present yet.\
You can also consult the data detail view for ongoing actions by clicking on the data in the overview. When clicking on an ongoing action itself, the data job details of the most recent created data job are shown.

### Folders

If you open a folder by clicking it, you can see the **folder details** link at the top right. This will open the details screen where you can consult the **folder size** and number of files in that folder, the owning project, **ongoing actions** and **folder id**. You can also download the folder and all contents here with the **download** button.

To help navigate between folders in flat view, you can use the "**path**' column in the data view which will open the folder containing the selected file in tree view. If you want to go further up the folder path, you can use the **folder structure above the file** **view**. If the path is not visible, you can add it with the three-columns symbol next to the filter symbol.

<figure><img src="/files/1UL4lNGXNsY1WtNGufDy" alt=""><figcaption></figcaption></figure>

### Searching for Data

To quickly find data, use the search dialog at the top right. Search is performed with automatic wildcards before and after the search text. You can use \* as additional wildcard, for example `b*n` will match `bunny` as it is interpreted as `*b*n*`.

You can search on the file name, the path (/folder/subfolder) and tags.

<figure><img src="/files/2tPLI825EPC5xau3EIDe" alt=""><figcaption></figcaption></figure>

### Secondary Data

When Secondary Data is added to a data record, those secondary data records are mounted in the same parent folder path as the primary data file when the primary data file is provided as an input to a pipeline. Secondary data is intended to work with the CWL [secondaryFiles](https://www.commonwl.org/v1.2/CommandLineTool.html#File) feature. This is commonly used with genomic data such as BAM files with companion BAM index files.

<figure><img src="/files/2EGaM1gWiPyM2jGAePKy" alt=""><figcaption></figcaption></figure>

***

## Hyperlinking to Data

You can create hyperlinks to data to quickly share it with the following syntax:

```
https://<ServerURL>/ica/link/project/<ProjectID>/data/<FolderID>
```

```
https://<ServerURL>/ica/link/project/<ProjectID>/analysis/<AnalysisID>
```

<table><thead><tr><th width="155">Variable</th><th>Location</th></tr></thead><tbody><tr><td><strong>ServerURL</strong></td><td>See browser address bar.</td></tr><tr><td><strong>projectID</strong></td><td>At YourProject > Details > URN > urn:ilmn:ica:project:<em><strong>ProjectID</strong></em>#MyProject</td></tr><tr><td><strong>FolderID</strong></td><td>At YourProject > Data > folder > folder details > <em><strong>ID</strong></em></td></tr><tr><td><strong>AnalysisID</strong></td><td>At YourProject > Flow > Analyses > YourAnalysis > <em><strong>ID</strong></em></td></tr></tbody></table>

{% hint style="info" %}
Normal permission checks still apply with these links. If you try to follow a link to data to which you do not have access, you will be returned to the main project screen or login screen, depending on your permissions.
{% endhint %}

***

## Exporting the Data List

You can export the list of data which you see in the overview as a CSV, JSON, or excel file.

1. Select one or more files to export at **Projects > your\_project > Data**.
2. Select **Export** at the bottom of the screen.
3. Choose between the following export options:
   * To export only the list of selected files, select the **Selected rows** as the Rows to export option. To export the list of all files on the page, select **Current page**.
   * To export only the columns which are currently shown in your view, select **Visible columns** as the Columns to export option, otherwise, choose **All columns**.
4. Select the export format. (CSV/JSON/Excel)

***

## Data Management

{% hint style="info" %}
To prevent cost issues, you can not perform actions such as copying and moving data which would write data to the workspace when the project billing mode is set to tenant and the owning tenant of the folder is not the current user's tenant.
{% endhint %}

### **Downloading Data**

**Single files** can be downloaded directly from within the UI.

* Select the checkbox next to the file which you want to download, followed by **Download > Browser Download > Download**.
* You can also download files from their details screen. Click on the file name and select Download at the bottom of the screen. Depending on the size of your file, it may take some time to load the file contents.

#### Schedule for Download

You can trigger an asynchronous download via service connector using the Schedule for Download button with one or more files selected.

1. Select a file or files to download.
2. Select **Download > Schedule download (for files or folders)**. This will display a list of all available connectors.
3. Select a connector and optionally, enter your email address if you want to be notified of download completion, and then select **Download**.

{% hint style="info" %}
If you do not have a connector, [create one](/project/p-connectivity/service-connector) and install it. You must then return to the file selection in step 1 to use it.
{% endhint %}

You can view the progress of the download or abort the scheduled download on the [Activity](/project/p-activity) page for the project.

### **Uploading Data**

Uploading data to the platform makes it available to analysis workflows and tools.

#### UI Upload

To upload data manually via the drag-and-drop interface in the platform UI, go to **Projects > your\_project > Data** and either

* Drag a file from your system into the *Choose a file or drag it here* box.
* Select the Choose a file or drag it here box, and then choose a file. Select Open to upload the file.

Your files are added to the Data page with status ***partial*** during upload and become ***available*** when upload completes.

{% hint style="info" %}
Do not close the Platform Core tab in your browser while data uploads.
{% endhint %}

<figure><img src="/files/MdHXN1KA7Bs7niWwAiwU" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Uploads via the UI are limited to 5TB and no more than 100 concurrent files at a time, but for practical and performance reasons, it is recommended to use the CLI or [Service connector](/project/p-connectivity/service-connector) when uploading large amounts of data.
{% endhint %}

#### Upload Data via CLI

For instructions on uploading/downloading data via CLI, see [CLI Data Transfer](/command-line-interface/cli-datatransfer).

***

### Copying Data

You can **copy** data from your project **to a different folder within the same project** or you can copy data **from another project to your current** project, provided you have the necessary access rights.

You can copy data **from a subfolder to a higher-level folder** to move data up one or more levels (folder/destination/source). You can not copy data from the source folder onto itself or onto a subfolder of the source folder as this would result in a loop.

{% hint style="info" %}
Copying large amounts of data can take considerable time. You can **monitor the progress** at **Projects > your\_project > Activity > Batch Jobs**.
{% endhint %}

#### Required Rights

The person copying the data must have the following rights:

<table><thead><tr><th width="184.33333333333331">Copy Data Rights</th><th>Source Project</th><th>Destination Project</th></tr></thead><tbody><tr><td>Within a project</td><td><ul><li>Contributor rights</li><li>Upload and Download rights</li></ul></td><td><ul><li>Contributor rights</li><li>Upload and Download rights</li></ul></td></tr><tr><td>Between different projects</td><td><ul><li>Download rights</li><li>Viewer rights</li></ul></td><td><ul><li>Upload rights</li><li>Contributor rights</li></ul></td></tr></tbody></table>

#### Restrictions

The following restrictions apply when copying data:

| Copy Data Restrictions     | Source Project                                                                                                         | Destination Project                                             |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- |
| Within a project           | <ul><li>No linked data</li><li>No partial data</li><li>No archived data</li></ul>                                      | <ul><li>No Linked data</li></ul>                                |
| Between different projects | <ul><li>Data sharing enabled</li><li>No partial data</li><li>No archived data</li><li>Within the same region</li></ul> | <ul><li>No linked data</li><li>Within the same region</li></ul> |

{% hint style="warning" %}
Data in the "Partial" or "Archived" state will be skipped during a copy job.
{% endhint %}

#### Copying Data

1. Go to the destination project for your data copy and proceed to **Projects > your\_project > Data > Manage > Copy From**.
2. Optionally, use the filters or search with the search box for the desired data.
3. Select the data (individual files or folders with data) you want to copy.
4. Select any meta data which you want to keep with the copied data (user tags, technical system tags or instrument information).
5. Select which action to take if the data already exists (overwrite existing data, don't copy or keep both the original and the new copy by appending a version number to the copied data).
6. Select **Copy** to copy the data to your project. You can see the progress in **Projects > your\_project > Activity > Batch Jobs** and if your browser permits it, a pop-up message will be displayed when the copy process completes.

<table data-view="cards"><thead><tr><th></th><th></th><th></th></tr></thead><tbody><tr><td><p><strong>Replace</strong></p><p>Overwrites the existing data. Folders will copy their data in an existing folder with existing files. Existing files will be replaced when a file with the same name is copied and new files will be added. The remaining files in the target folder will remain unchanged.</p></td><td></td><td></td></tr><tr><td><p><strong>Don't copy</strong></p><p>The original files are kept. If you selected a folder, <em>files that do not yet exist in the destination folder are added to</em> it. Files that already exist at the destination are not copied over and the originals are kept.</p></td><td></td><td></td></tr><tr><td><p><strong>Keep both</strong></p><p>Files have a number appended to them if they already exist. If you copy folders, the folders are merged, with new files added to the destination folder and original files kept. New files with the same name get copied over into the folder with a number appended.</p></td><td></td><td></td></tr></tbody></table>

{% hint style="info" %}
There is a difference in copy type behavior between copying files and folders. The behavior is designed for files and it is **best practice to not copy folders if there already is a folder with the same name in the destination location**.
{% endhint %}

#### Copy Status

* INITIALIZED
* WAITING\_FOR\_RESOURCES
* RUNNING
* STOPPED - When choosing to stop the batch job.
* SUCCEEDED - All files and folders are copied.
* PARTIALLY\_SUCCEEDED - Some files and folders could be copied, but not all. Partially succeeded will typically occur when files were being modified or unavailable while the copy process was running.
* FAILED - None of the files and folders could be copied.

To see the ongoing actions on data in the data overview (**Projects > your\_project > Data**), you can add the **ongoing actions** column from the column list with the three column symbol at the top right, next to the filter funnel. You can also consult the data detail view for ongoing actions by clicking on the data in the overview.

{% hint style="info" %}
Notes on copying data

* Copying data comes with an additional storage cost as it will create a copy of the data.
* Copying data from your own S3 storage requires additional configuration. See [Connect AWS S3 Bucket](/home/h-storage/s-awss3) and [SSE-KMS Encryption](/home/h-storage/s-awss3/s-sse-kms).
* On the command-line interface, the command to copy data is `icav2 projectdata copy`.
* Before copy and move operations are executed **on your own S3 storage**, a test is performed to verify the necessary operational rights. This can result in temporary test files remaining (for example when [IAM policy](/home/h-storage/s-awss3/iam-role-method#id-2-create-data-access-permission-aws-iam-policy) is not correctly set up for a versioned bucket). These files can safely be manually deleted from your S3 console.
  {% endhint %}

***

### Moving Data

You can move data within a project or between different projects to which you have access. If your browser allows notifications, a pop-up will appear when the move is completed.

* **Move From** is used when you are **in the destination** location.
* **Move To** is used when you are **in the source** location.

Before moving the data, pre-checks are performed to verify that the data can be moved and no currently running operations are being performed on the folder. Conflicting jobs and missing permissions will be reported.

Once the move has started, no other operation must be performed on the data being moved to avoid potential data loss or duplication. When modifying data at the source or destination during a move process, incomplete data transfers may occur with duplicate folders and files with different identifiers.. You can manually transfer any remaining data and delete duplicate files and folders afterward.

<details>

<summary>Move Synchronization issues</summary>

Changes to the date during move may cause the destination data to be unsynchronized between the object store (S3) and Platform Core. To address this, create a folder session on the destination directory's parent folder by using the following API steps: [Create Folder Session](https://ica.illumina.com/ica/api/swagger/index.html#/Project%20Data/createFolderUploadSession) and [Complete Folder Session](https://ica.illumina.com/ica/api/swagger/index.html#/Project%20Data/completeFolderUploadSession). Ensure that the move job is aborted before making the create and complete requests for the folder session.

</details>

{% hint style="warning" %}
Move jobs will fail if any data being moved is in the *Partial* or *Archived* state.
{% endhint %}

#### Required Rights

There are a number of rights and restrictions related to data move as this will delete the data in the source location.

| Move Data Rights           | Source Project                                               | Destination Project                                   |
| -------------------------- | ------------------------------------------------------------ | ----------------------------------------------------- |
| Within a project           | <ul><li>Contributor rights</li></ul>                         | <ul><li>Contributor rights</li></ul>                  |
| Between different projects | <ul><li>Download rights</li><li>Contributor rights</li></ul> | <ul><li>Upload rights</li><li>Viewer rights</li></ul> |

#### Restrictions

<table><thead><tr><th>Move Data Restrictions</th><th width="286">Source Project</th><th>Destination Project</th></tr></thead><tbody><tr><td>Within a project</td><td><ul><li>No linked data</li><li>No partial data</li><li>No archived data</li></ul></td><td><ul><li>No Linked data</li></ul></td></tr><tr><td>Between different projects</td><td><ul><li>Data sharing enabled</li><li>Data owned by user's tenant</li><li>No linked data</li><li>No partial data</li><li>No archived data</li><li>No externally managed projects</li><li>Within the same region</li></ul></td><td><ul><li>No linked data</li><li>Within same region</li></ul></td></tr></tbody></table>

#### Moving Data Constraints

* **1000 Maximum Items:** Up to 1000 items per move. Items include files and folders. Folders with subfolders and subfiles still count as one item.
* **Naming Conflicts:** Cannot move to a destination with existing files/folders of the same name.
* **Linked Data Restrictions:** Cannot move linked data move data to linked data.
* **Self Move:** Folders cannot be moved to themselves.
* **In-Transit Data:** Cannot move data that is being moved.
* **Region Restrictions:** No cross-region moves allowed.
* **Project Constraints:** No moves from externally-managed projects or externally-managed data.
* **Status Requirement:** Data must be in status available.
* **Ownership:** Data must be owned by the user's tenant for cross-project moves.
* **Destination Default:** If no target folder is selected, data moves to the root folder of the target project.

#### Move Data From

Move Data From is used when you are in the **destination** location.

1. Navigate to **Projects > your\_project > Data > your\_destination\_location > Manage > Move From**.
2. Select the files and folders which you want to move.
3. Select the **Move** button.

{% hint style="info" %}
Moving large amounts of data can take considerable time. You can **monitor the progress** at **Projects > your\_project > Activity > Batch Jobs**.
{% endhint %}

#### Move Data To

Move Data To is used when you are in the **source** location. You will need to select the data you want to move from to current location and the destination to move it to.

1. Navigate to **Projects > your\_project > Data > your\_source\_location**.
2. Select the files and folders which you want to move.
3. Select to **Projects > your\_project > Data > your\_source\_location > Manage > Move To**.
4. Select your target project and location.\
   You can create a new folder to move data to by filling in the "New folder name (optional)" field. This does NOT rename an existing folder. **To rename a folder**, you will need to create a new folder with the desired name, move the contents from the original folder into the new one, and then delete the original folder.
5. Select the **Move** button.

{% hint style="info" %}
Moving large amounts of data can take considerable time. You can **monitor the progress** at **Projects > your\_project > Activity > Batch Jobs**.
{% endhint %}

#### Move Status

* INITIALIZED
* WAITING\_FOR\_RESOURCES
* RUNNING
* STOPPED - When choosing to stop the batch job.
* SUCCEEDED - All files and folders are moved.
* PARTIALLY\_SUCCEEDED - Some files and folders could be moved, but not all. Partially succeeded will typically occur when files were being modified or unavailable while the move process was running.
* FAILED - None of the files and folders could be moved.

To see the ongoing actions on data in the data overview (**Projects > your\_project > Data**), add the **ongoing actions** column from the column list with the three column symbol at the top right, next to the filter funnel. You can also consult the data detail view for ongoing actions by clicking on the data in the overview.

{% hint style="info" %}
If you are only able to select your source project as the target data project, this may indicate that **data sharing** (**Projects > your\_project > Project Settings > Details > Data Sharing**) is not enabled for your project or that you do not have have upload rights in other projects.
{% endhint %}

{% hint style="info" %}
Before copy and move operations are executed **on your own S3 storage**, a test is performed to verify the necessary operational rights. This can result in temporary test files remaining (for example when [IAM policy](/home/h-storage/s-awss3/iam-role-method#id-2-create-data-access-permission-aws-iam-policy) is not correctly set up for a versioned bucket). These files can safely be manually deleted from your S3 console.
{% endhint %}

***

### Deleting, Archiving and Unarchiving

To manually archive or delete files:

1. Select the checkbox next to the file or files to delete or archive.
2. Select Manage, and then select one of the following options:
   * **Archive** — Move the file or files to long-term storage (event code ICA\_DATA\_110).
   * **Unarchive** — Return the file or files from long-term storage. Unarchiving can take up to 48 hours, regardless of file size. Unarchived files can be used in analysis (event code ICA\_DATA\_114).
   * **Delete** — Remove the file completely (event code ICA\_DATA\_106).

When attempting concurrent archiving or unarchiving of the same file, a message will inform you to wait for the currently running (un)archiving to finish first.

To archive or delete files programmatically, you can use Platform Core's API endpoints:

1. [GET](https://ica.illumina.com/ica/api/swagger/index.html#/Project%20Data/getProjectData) the file's information.
2. Modify the dates of the file to be deleted/archived.
3. [PUT](https://ica.illumina.com/ica/api/swagger/index.html#/Project%20Data/updateProjectData) the updated information back in Platform Core.

<details>

<summary>Python Example</summary>

The Python snippet below exemplifies the approach: it sets (or updates if set already) the time to be archived for a specific file:

```python
import requests
import json

from config import PROJECT_ID, DATA_ID, API_KEY

url_get="https://ica.illumina.com/ica/rest/api/projects/" + PROJECT_ID + "/data/" + DATA_ID

# set the API get headers
headers = {
            'X-API-Key': API_KEY,
            'accept': 'application/vnd.illumina.v3+json'
            }

# set the API put headers
headers_put = {
            'X-API-Key': API_KEY,
            'accept': 'application/vnd.illumina.v3+json',
            'Content-Type': 'application/vnd.illumina.v3+json'
            }

# Helper function to insert willBeArchivedAt after field named 'region'
def insert_after_region(details_dict, timestamp):
    new_dict = {}
    for k, v in details_dict.items():
        new_dict[k] = v
        if k == 'region':
            new_dict['willBeArchivedAt'] = timestamp
    if 'willBeArchivedAt' in details_dict:
        new_dict['willBeArchivedAt'] = timestamp
    return new_dict

# 1. Make the GET request
response = requests.get(url_get, headers=headers)
response_data = response.json()

# 2. Modify the JSON data
timestamp = "2024-01-26T12:00:04Z"  # Replace with the provided timestamp
response_data['data']['details'] = insert_after_region(response_data['data']['details'], timestamp)

# 3. Make the PUT request
put_response = requests.put(url_get, data=json.dumps(response_data), headers=headers_put)
print(put_response.status_code)
```

To delete a file at specific timepoint, the key 'willBeDeletedAt' should be added or changed using the API call. If running in the terminal, a successful run will finish with the message ‘200’. In the Platform Core UI, you can check the details of the file to see the updated values for ‘Time To Be Archived’ (willBeArchivedAt) or ‘Time To Be Deleted’ (willBeDeletedAt), as shown in the screenshot.

<img src="/files/fnGZtds2G08qOx30ooM0" alt="" data-size="original">

</details>

***

### Linking and Unlinking

Data linking creates a **dynamic** **read-only view** to the source data. You can use data linking to get access to data without running the risk of modifying the source material and to share data between projects. Linking ensures **changes to the source data are immediately visible** and **no additional storage** is required. You can recognise linked data by the **green color** and see the owning project as part of the details.

Since this is read-only access, you cannot perform actions such as deleting, adding, moving or (un)archiving on linked data as these actions require write access.

{% hint style="info" %}
**Linking data is only possible from the root folder of your destination project.** The action is disabled in project subfolders.

Linking a parent folder after linking a file or subfolder will unlink the file or subfolder and link the parent folder. So *root\linked\_subfolder* will become *root\linked\_parentfolder\linked\_subfolder*.
{% endhint %}

{% hint style="info" %}
Initial linking can take considerable time when there is a large amount of source data. However, once the initial link is made, updates to the source data will be instantaneous. You can monitor the progress at **Projects > your\_project > activity > Batch Jobs**.
{% endhint %}

<details>

<summary><strong>Migrating snapshot linked data. (linked before Platform Core release v.2.29)</strong></summary>

Before Platform Core version v.2.29, when data was linked, a snapshot was created of the file and folder structure. These links created a read-only view of the data as it was at the time of linking, but did not propagate changes to the file and folder structure. If you want to use the advantages of the new way of linking with dynamic updates, unlink the data and relink it. Since snapshot linking has been deprecated, **all new data linking done in Platform Core v.2.29 or later has dynamic content updates.**

</details>

#### Linking data from another project.

1. Select **Projects > your\_project > Data > Manage**, and then select **Link**.
2. To view data by project, select the **funnel symbol**, and then select **Owning Project**. If you know to which project the data is linked to, you can choose to filter on linked projects. If you click on a folder, the folder will open so you can access the files, if you click on a file, the file details will be opened.
3. Select the checkbox next to the file or files to add.
4. Select **Link**.

Your files will be added and visible in the Data page.

<details>

<summary>Display Owning Project</summary>

if you have selected multiple owning projects, you can add the owning project column to see which project owns the data.

1. At the top of the screen, next to the filer icon, select the **three columns**.
2. The Add/remove columns tab will appear.
3. Choose **Owning Project** (or Linked Projects)

   <figure><img src="/files/vrhX4dzlhEoFVpiH3vb8" alt=""><figcaption><p>Owning Project Filter</p></figcaption></figure>

</details>

#### Linking Folders

If you link a **folder** instead of individual files, a warning is displayed indicating that, depending on the size of the folder, linking may take considerable time. The linking process will run in the background and the progress can be monitored on the **Projects > your\_project > activity > Batch Jobs** screen. From here you can see more details such as how many files have already been linked, by clicking the batch job.

#### Unlinking Project Data

To unlink the data, go to the root level of your project and select the linked folder or, if you have linked individual files separately, then you can select those linked files (limited to 100 at a time) and select **Manage > Unlink**. The progress can be monitored at **Projects > your\_project > Activity > Batch Jobs**.


# Non-Indexed Folders

Non-indexed folders (<img src="/files/sObxgZsxRbzCsbftmMFG" alt="" data-size="line">) are designed for optimal performance in situations where no file actions are needed. They serve as fast storage in situations like temporary analysis file storage where you don't need access or searches via the GUI to individual files or subfolders within the folder. Think of a non-indexed folder as a **data container**. You can access the container which contains all the data, but you can **not access the individual data files within the container from the GUI**. As non-indexed folders contain data, they count towards your total project storage.

You can see the size of a non-indexed folder as part of the data details screen (**Projects > your\_project > Data > your\_non-indexed\_data > Data details tab**) and in the data view (**Projects > your\_project > Data)**.

{% hint style="info" %}
There can be a noticeable delay before the size of a non-indexed folder is updated after changes because of how the data is handled.
{% endhint %}

The GUI considers non-indexed folders as a single object. You can access the contents from a non-indexed folder

* as Analysis input/output
* in Bench
* via the API

<table><thead><tr><th width="169">Action</th><th width="100">Allowed</th><th>Details</th></tr></thead><tbody><tr><td>Creation</td><td>Yes</td><td>You can create non-indexed folders at <strong>Projects > your_project > Data > Manage > Create non-indexed folder</strong>. or with the <code>/api​/projects​/{projectId}​/data:createNonIndexedFolder</code> <a href="https://ica.illumina.com/ica/api/swagger/index.html">endpoint</a></td></tr><tr><td>Deletion</td><td>Yes</td><td>You can delete non-indexed folders by selecting them at <strong>Projects > your_project > Data > select the folder > Manage > Delete.</strong><br>or with the <code>/api​/projects​/{projectId}​/data/{dataId}:delete</code> endpoint</td></tr><tr><td>Uploading Data</td><td>API<br>Bench<br>Analysis</td><td>Use non-indexed folders as normal folders for Analysis runs and bench. Different methods are available with the API such as creating temporary credentials to upload data to S3 or using <code>/api/projects/{projectId}/data:createFileWithUploadUrl</code></td></tr><tr><td>Downloading Data</td><td>Yes</td><td>Use non-indexed folders as normal folders for Analysis runs and bench. Use temporary credentials to list and download data with the API.</td></tr><tr><td>Analysis Input/Output</td><td>Yes</td><td>Non-indexed files can be used as input for an analysis and the non-indexed folder can be used as output location. You will not be able to view the contents of the input and output in the analysis details screen.</td></tr><tr><td>Bench</td><td>Yes</td><td>Non-indexed folders can be used in Bench and the output from Bench can be written to non-indexed folders. Non-indexed folders are accessible across Bench workspaces within a project.</td></tr><tr><td>Viewing</td><td>No</td><td>The folder is a single object, you can not view the contents.</td></tr><tr><td>Linking</td><td>Yes</td><td>You cannot see non-indexed folder contents.</td></tr><tr><td>Copying</td><td>No</td><td>Prohibited to prevent storage issues.</td></tr><tr><td>Moving</td><td>No</td><td>Prohibited to prevent storage issues.</td></tr><tr><td>Managing tags</td><td>No</td><td>You cannot see non-indexed folder contents.</td></tr><tr><td>Managing format</td><td>No</td><td>You cannot see non-indexed folder contents.</td></tr><tr><td>Use as Reference Data</td><td>No</td><td>You cannot see non-indexed folder contents.</td></tr></tbody></table>


# Data Integrity

You can verify the integrity of the data by comparing the hash which is usually ([with some exceptions](https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity.html#checking-object-integrity-md5)) an MD5 (Message Digest Algorithm 5) checksum. This is a common cryptographic hash function that generates a fixed-size, 128-bit hash value from any input data. This hash value is unique to the content of the data, meaning even a slight change in the data will result in a significantly different MD5 checksum. AWS S3 calculates this checksum when data is uploaded and stores it in the ETag (Entity tag).

For files smaller than 16 MB, you can directly retrieve the MD5 checksum using our [API](https://ica.illumina.com/ica/api/swagger/index.html) endpoints. Make an API GET call to the `https://ica.illumina.com/ica/rest/api/projects/{projectId}/data/{dataId}` endpoint specifying the data Id you want to check and the corresponding project ID. The response you receive will be in JSON format, containing various file metadata. Within the JSON response, look for the `objectETag` field. This value is the MD5 checksum for the file you have queried. You can compare this checksum with the one you compute locally to ensure file integrity.

This ETag does not change and can be used as a file integrity check even when that file is archived, unarchived and/or copied to another location. Changes to the metadata have no impact on the ETag

For larger files, the process is different due to computation limitations. In these cases, we recommend using a dedicated pipeline on our platform to explicitly calculate the MD5 checksum. Below you can find both a main.nf file and the corresponding XML for a possible Nextflow pipeline to calculate the MD5 checksum for FASTQ files.

<pre><code><strong>nextflow.enable.dsl = 2
</strong>

process md5sum {
    
    container "public.ecr.aws/lts/ubuntu:22.04"
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'
    
    input:
        file txt

    output:
        stdout emit: result
        path '*', emit: output

    publishDir "out", mode: 'symlink'

    script:
        txt_file_name = txt.getName()
        id = txt_file_name.takeWhile { it != '.'}

        """
        set -ex
        echo "File: $txt_file_name"
        echo "Sample: $id"
        md5sum ${txt} > ${id}_md5.txt
        """
    }

workflow {
    txt_ch = Channel.fromPath(params.in)
    txt_ch.view()
    md5sum(txt_ch).result.view()
}
</code></pre>

{% code overflow="wrap" fullWidth="false" %}

```
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition">
    <pd:dataInputs>
        <pd:dataInput code="in" format="FASTQ" type="FILE" required="true" multiValue="true">
            <pd:label>Input</pd:label>
            <pd:description>FASTQ files input</pd:description>
        </pd:dataInput>
    </pd:dataInputs>
    <pd:steps/>
</pd:pipeline>
```

{% endcode %}


# Storage Cost Managment

Since there is a storage cost associated with the data in your projects, it is good practice to regularly check how much cost is being generated by your projects and evaluate which data can be removed from cloud storage. The instructions provided here will help you quickly determine which data is generating the highest storage costs.

### Monitoring Storage Cost

To see how much storage costs are currently being generated for your tenant, you can look at the [usage explorer](https://help.connected.illumina.com/account-management/usage-explorer) at <https://platform.illumina.com/usage/> or from within Platform Core, navigate to the 9-dot symbol (<img src="/files/Uy4JoigujLLGJC1ZK3lk" alt="" data-size="line">) in the top right next to your name and choose the usage explorer from the menu.

From the usage explorer overview screen, you can see below the graphical representation which **projects** are incurring the highest storage costs.

<figure><img src="/files/njxKIyeUIbMtw0nW1cNG" alt=""><figcaption></figcaption></figure>

### Project Files in Platform Core

When you have determined which projects are incurring the largest storage costs, you can find out which files within that project are taking up the most space. To find the largest files in your project,

1. Go to **Projects > your\_project > Data** and switch to list view with the (<img src="/files/1aSZTyHL7AbUMWBrxknY" alt="" data-size="line">) icon left above your files.
2. Use the column icon (<img src="/files/S2H9ZCGfWjlr5YQPLr8D" alt="" data-size="line">) top right to add the size column to your view. You can drag the size column to the desired position in your list view or use the move left and use right options which appear when you select the three vertical dots.
3. Select Sort descending to show the largest files first.

<figure><img src="/files/uCNJwgY0YsKZS3GH1tJz" alt="" width="375"><figcaption></figcaption></figure>

4. Once you have the list sorted like this, you can evaluate if those large files are still needed, if the can be sent to [archive](/project/p-data#archiving-and-deleting-files) (**manage > archive**) or if they can be deleted (**manage > delete**).


# Data - Troubleshooting

### Path too long

If you get an error "Unable to generate credentials from the objectstore as the requested path is too long." from AWS when requesting temporary credentials, then the path should be shortened.

You can truncate the sample name and user reference or use advanced output mapping in the API which avoids generating the long folders and creates output in the targetPath-defined location.

```
"analysisOutput": [
{
"sourcePath": "out",
"type": "FOLDER",
"targetProjectId": "enter_your_target_project_id",
"targetPath": "/enter_your_target_folder/"
}
]
```

### Files remaining on own S3 Storage

Before copy and move operations are executed **on your own S3 storage**, a test is performed to verify the necessary operational rights. If an issue is encountered, the copy/move action is aborted and the test files can be left behind due to missing rights. These files can safely be manually deleted from your S3 console.


# Samples

You can use samples to group information related to a sample, including input files, output files, and analyses. You can consider samples as creating a binder to collect related information. When you then link that sample to another project, you bring over the empty binder, but not the files contained in it. In that project, you can then add your own data to it.

You can search for samples (excluding their metadata) with the Search button at the top right.

## Add New Sample

To add a new sample, do as follows.

1. Select **Projects > your\_project > Samples**.
2. To add a new sample, select **+ Create**, and then enter a unique name and description for the sample.
3. To add files to the sample, see [adding files to a sample](#add-files-to-samples)

Your sample is added to the Samples page. To view information on the sample, select the sample name in the overview.

## Add Files to Samples

You can add files to a sample after creating the sample. Any files that are not currently linked to a sample are listed on the Unlinked Files tab.

To add an unlinked file to a sample, do as follows.

1. Go to **Projects > your\_project > Samples > Unlinked files tab**.
2. Select a file or files, and then select one of the following options:
   * **Create sample** — Create a new sample that includes the selected files.
   * **Link to sample** — Select an existing sample in the project to which you link the file.

Alternatively, you can add unlinked files from the sample details.

1. Going to **Projects > your\_project > Samples > your\_sample**.
2. Select your sample to open the details.
3. Go to the **Data tab** and select **link**.
4. If the data is not in your project, you will need to add it in **Projects > your\_project > Data**

Data can only be linked to a single sample, so once you have linked data to a sample, it will no longer appear in the list of data to choose form.

## Unlink Files from Samples

To remove files from samples,

1. Go to **Projects > your\_project > Samples > your\_sample > Data**
2. Select the files you want to remove.
3. Select **Unlink**.

{% hint style="info" %}
If your selection contains both unlinkable samples and non-unlinkable samples (for example linked via a linked bundle), then this will be indicated in the confirmation dialog and only the unlinkable samples will be unlinked.
{% endhint %}

## Link Samples to Project

A Sample can be linked to a project from a separate project to make it available in read-only capacity.

1. Navigate to the Samples view in the Project
2. Click the **Link** button
3. Select the Sample(s) to link to the project
4. Click the **Link** button

{% hint style="warning" %}
Data linked to Samples is not automatically linked to the project. The data must be linked separately from the Data view. Samples also must be available in a complete state in order to be linked.
{% endhint %}

## Delete Samples

If you want to remove a sample, select it and use the delete option from the top navigation row. You will be presented a choice of how to handle the data in the sample.

* Unlink all data without deleting it.
* Delete input data and unlink other data.
* Delete all data.


# Activity

The Activity view (**Projects > your\_project > Activity**) shows the status and history of long-running activities including Data Transfers, Base Jobs, Base Activity, Bench Activity and Batch Jobs.

## Data Transfers

The Data Transfers tab shows the status of data uploads and downloads. You can sort, search and filter on various criteria and export the information. **Show ongoing transfers** (top right) allows you to filter out the completed and failed transfers to focus on current activity.

Transfers with a yellow background indicate that [service connector](/project/p-connectivity/service-connector) rules have been modified in ways that prevent planned files from being uploaded. Please verify your service connectors to resolve this.

## Base Jobs

The Base Jobs tab gives an overview of all the actions related to a table or a query that have run or are running (e.g., Copy table, export table, Select \* from table, etc.) If a job is still running, you can **abort** it from this screen

The jobs are shown with their:

* **Creation time**: When did the job start
* **Description**: The query or the performed action with some extra information
* **Type**: Which action was taken
* **Status**: Failed or succeeded
* **Duration**: How long the job took
* **Billed bytes**: The used bytes that need to be paid for

Failed jobs provide information on why the job failed. Click the job to see more details. If a job is retried, the original failed job will still remain visible here.

## Base Activity

The Base Activity tab gives an overview of previous results (e.g., Executed query, Succeeded Exporting table, Created table, etc.) Collecting this information **may take some time**. For performance reasons, only the activity of the **last month** (rolling window) with a **limit of 1000** records is shown.

You can use the **Export** function at the bottom of the screen to download the **current page or selected rows** in Excel, JSV or JSON format.

To get the **data for the last year** without limit on the number of records, use the **export to project file** function at the top of the screen. This will export the data in JSON format as either a single file or split over multiple files according to your desired file size. No activity data is retained for more than one year.

The activities are shown with:

* **Start Time**: The moment the action was started
* **Query**: The SQL expression.
* **Status**: Failed or succeeded
* **Duration**: How long the job took
* **User**: The user that requested the action
* **Size**: For SELECT queries, the size of the query results is shown. Queries resulting in less than 100Kb of data will be shown with a size of <100K

## Bench Activity

The Bench Activity tab shows the actions taken on Bench Workspaces in the project.

The activities are shown with:

* **Workspace**: Workspace where the activity took place
* **Date**: Date and time of the activity
* **User**: User who performed the activity
* **Action**: Which activity was performed

## Batch Jobs

The Batch Jobs tab allows users to monitor progress of Batch Jobs in the project. It lists Data Downloads, Sample Creation (double-click entries for details) and Data Linking (double-click entries for details). The (ongoing) Batch Job details are updated each time they are (re)opened, or when the refresh button is selected at the bottom of the details screen. Batch jobs which have a final state such as Failed or Succeeded are removed from the activity list after 7 days.

Which batch jobs are visible depends on the user role.

| Project Creator | <p>Project Collaborator<br>same tenant</p> | <p>Project Collaborator<br>different tenant</p> |
| --------------- | ------------------------------------------ | ----------------------------------------------- |
| All batch jobs  | All batch jobs                             | Only batch jobs of own tenant                   |


# Flow

Flow provides tooling for building and running secondary analysis pipelines. The platform supports analysis workflows constructed using Common Workflow Language (CWL) and Nextflow. Each step of an analysis pipeline executes a containerized application using inputs passed into the pipeline or output from previous steps.

You can configure the following components in Illumina Connected Analytics Flow:

* Reference Data — Reference Data for Graphical CWL flows. See [Reference Data](/project/p-flow/f-referencedata).
* Pipelines — One or more tools configured to process input data and generate output files. See [Pipelines](/project/p-flow/f-pipelines).
* Analyses — Launched instance of a pipeline with selected input data. See [Analyses](/project/p-flow/f-analyses).


# Reference Data

**Reference Data** are reference genome sets which you use to help look for deviations and to compare your data against.

## Creating Reference Data

**Reference data** properties are located at the **main navigation level** and consist of the following free text fields.

* Types
* Species
* Reference Sets

Once these are configured,

1. Go to your data in **Projects > your\_project > Data**.
2. Select the data you want to use as reference data and **Manage > Use as reference data**.
3. Fill out the configuration screen

<figure><img src="/files/ItIYZPscsTeFG0OftRmF" alt="" width="375"><figcaption></figcaption></figure>

You can see the result at the main navigation level > **Reference Data** (outside of projects) or in **Projects > your\_project > Flow > Reference Data**.

## Linking Reference Data to your Project

To use a reference set from within a project, you have first to add it. Select **Projects > your\_project > Flow > Reference Data > Link**. Then select a reference set to add to your project.

{% hint style="info" %}
Reference sets are only supported in Graphical CWL pipelines.
{% endhint %}

## Copying Reference Data to other Regions

1. Navigate to **Reference Data** (Not from within a project, but outside of project context, so at the main navigation level).
2. Select the data set(s) you wish to add to another region and select **Copy to another project**.
3. Select a project located in the region where you want to add your reference data.
4. You can check in which region(s) Reference data is present by opening the Reference set and viewing **Copy Details**.
5. Allow a few minutes for new copies to become available before use.

{% hint style="info" %}
You only need one copy of each reference data set per region. Adding Reference Data sets to additional projects set in the same region does not result in extra copies, but creates links instead. This is done from inside the project at **Projects > \<your\_project> > Flow > Reference Data > Manage > Add to project**.
{% endhint %}

## Creating a Pipeline with Reference Data

To create a pipeline with a reference data, use the [CWL - graphical](/project/p-flow/f-pipelines#create-a-pipeline) mode. **Projects > your\_project > Flow > Pipelines > +Create > CWL Graphical**. Use the reference data icon instead of regular input icon. On the right hand side use the *Reference files* submenu to specify the name, the format, and the filters. You can specify the options for an end-user to choose from and a default selection. You can select more than 1 file, but you can only select 1 at a time (so, repeat process to select multiple reference files). If you only select 1 reference file, that file will be the only one users can use with your pipeline. In the screenshot a reference data with two options is presented.

{% hint style="info" %}
Safari is not supported for graphical CWL data.
{% endhint %}

![Two options for a reference file](/files/UngO8ipnkDomoqnja1Wn)

If your pipeline was built to give users the option of choosing among multiple input reference files, they will see the option to select among the reference files you configured, under Settings. After clicking the magnifying glass icon the user can select from provided options.


# Pipelines

A Pipeline is a series of Tools with connected inputs and outputs configured to execute in a specific order.

## Linking Existing Pipelines

Linking a pipeline (**Projects > your\_project > Flow > Pipelines > Link**) adds that pipeline to your project. This is not as a copy, but as the actual pipeline, so any changes to the pipeline are atomatically propagated to and from any project which has this pipeline linked.

You can link a pipeline if it is not already linked to your project and it is from your tenant or available in your [bundle](/home/h-bundles) or activation code.

{% hint style="info" %}
**Activation codes** are tokens which allow you to run your analyses and are used for accounting and allocating the appropriate resources. Platform Core will automatically determine the best matching activation code, but this can be overwritten if needed.
{% endhint %}

If you **unlink a pipeline** it removes the pipeline from your project, but it remains part of the list of pipelines of your tenant, so it can be linked to other projects later on.

{% hint style="info" %}
There is no way to permanently delete a pipeline.
{% endhint %}

***

## Create a Pipeline

Pipelines are created and stored within projects.

1. Navigate to **Projects > your\_project > Flow > Pipelines > +Create**.
2. Select **Nextflow** (XML / JSON / Git) , **CWL Graphical** or **CWL code** (XML / JSON / Git) to create a new Pipeline.
3. Configure pipeline settings in the pipeline property tabs.
4. When creating a graphical CWL pipeline, drag connectors to link tools to input and output files in the canvas. Required tool inputs are indicated by a yellow connector.
5. Select Save.

{% hint style="warning" %}
Pipelines use the latest tool definition when the pipeline was last saved. Tool changes do not automatically propagate to the pipeline. In order to update the pipeline with the latest tool changes, edit the pipeline definition by removing the tool and re-adding it back to the pipeline.
{% endhint %}

{% hint style="info" %}
Individual Pipeline files are limited to 20 Megabytes. If you need to add more than this, split your content over multiple files.
{% endhint %}

### Pipeline Statuses

For pipeline authors sharing and distributing their pipelines, the **draft**, **released**, **deprecated**, and **archived** statuses provide a structured framework for managing pipeline availability, user communication, and transition planning. To change the pipeline status, select it at **Projects > your\_project > Pipelines > your\_pipeline > change status.**

{% hint style="info" %}
You can edit pipelines while they are in *Draft* status. Once they move away from draft, pipelines can no longer be edited. Pipelines can be cloned (top right in the details view) to create a new editable version.
{% endhint %}

<table data-view="cards"><thead><tr><th>Status</th><th>Purpose</th><th>Best Practice</th></tr></thead><tbody><tr><td><strong>Draft</strong></td><td>Use the draft status while <strong>developing or testing</strong> a pipeline version internally.</td><td>Only share draft pipelines with collaborators who are actively involved in development.</td></tr><tr><td><strong>Released</strong></td><td>The released status signals that a pipeline is stable and ready <strong>for general use</strong>.</td><td>Share your pipeline when it is ready for broad use. Ensure users have access to current documentation and know where to find support or updates. Releasing a pipeline is only possible if all tools of that pipeline must be in released status.</td></tr><tr><td><strong>Deprecated</strong></td><td>Deprecation is used when a pipeline version is <strong>scheduled for retirement or replacement</strong>. Deprecated pipelines <strong>can not be linked</strong> to bundles, but will not be unlinked from existing bundles. Users who already have access will <strong>still be able to start analyses</strong>. You can add a message (max 256 chars) when deprecating pipelines.</td><td>Deprecate in advance of archiving a pipeline, making sure the new pipeline is available in the same bundle as the deprecated pipeline. This will allow the pipeline author to link the new or alternative pipeline in the deprecation message field.</td></tr><tr><td><strong>Archived</strong></td><td>Archiving a pipeline version removes it from active use; users <strong>can no longer launch</strong> <strong>analyses</strong>. Archived pipelines can not be linked to bundles, but are not automatically unlinked from bundles or projects. You can add a message (max 256 chars) when archiving pipelines.</td><td>Warn users in advance: Deprecate the pipeline before archiving to allow existing users time to transition. Use the archive message to point users to the new or alternative pipeline</td></tr></tbody></table>

***

## Pipeline Properties

The following sections describe the properties that can be configured in each tab of the pipeline editor.

Depending on how you design the pipeline, the displayed tabs differ between the graphical and code definitions. For **CWL** you have a choice on how to define the pipeline, **Nextflow** is always defined in code mode.

<table data-view="cards"><thead><tr><th></th><th></th><th></th></tr></thead><tbody><tr><td><strong>CWL Graphical</strong></td><td><ul><li>Details</li><li>Documentation</li><li>Definition</li><li>Analysis Report</li><li>Metadata Model</li><li>Report</li></ul></td><td></td></tr><tr><td><strong>CWL Code</strong></td><td><ul><li>Details</li><li>Documentation</li><li>Inputform files (JSON) or XML Configuration (XML)</li><li>CWL Files</li><li>Metadata Model</li><li>Report</li></ul></td><td></td></tr><tr><td><strong>Nextflow Code</strong></td><td><ul><li>Details</li><li>Documentation</li><li>Inputform Files (JSON) or XML Configuration (XML)</li><li>Nextflow files</li><li>Metadata Model</li><li>Report</li></ul></td><td></td></tr></tbody></table>

Any additional source files related to your pipeline will be displayed here in alphabetical order.

See the following pages for language-specific details for defining pipelines:

* [Nextflow](/project/p-flow/f-pipelines/pi-nextflow)
* [CWL](/project/p-flow/f-pipelines/pi-cwl)

***

### Details

The details tab provides options for configuring basic information about the pipeline.

<table><thead><tr><th width="216">Field</th><th>Entry</th></tr></thead><tbody><tr><td>Code (pipeline name)</td><td>The name of the pipeline. The name must be unique within the tenant, including linked and unlinked pipelines.</td></tr><tr><td>Nextflow Version</td><td>User selectable Nextflow version available only for Nextflow pipelines</td></tr><tr><td>Description</td><td>A short description of the pipeline.</td></tr><tr><td>Status</td><td>The <a href="#pipeline-statuses">release status</a> of the pipeline.</td></tr><tr><td>Proprietary</td><td>Hide the pipeline scripts and details from users who do not belong to the tenant who owns the pipeline. This also prevents cloning the pipeline.</td></tr><tr><td>Storage size</td><td>User selectable <a href="/pages/rewnkXVCIALBvAXMmNdn#data-storage">storage size</a> for running the pipeline. This must be large enough to run the pipeline, but setting it too large incurs unnecessary costs.</td></tr><tr><td>Links</td><td>External reference links. (max 100 chars as name and 2048 chars as link)</td></tr></tbody></table>

The following information becomes visible when viewing the pipeline details.

<table><thead><tr><th width="217">Field</th><th>Entry</th></tr></thead><tbody><tr><td>ID</td><td>Unique Identifier of the pipeline.</td></tr><tr><td>URN</td><td>Identification of the pipeline in Uniform Resource Name</td></tr></tbody></table>

The **clone** action will be shown in the pipeline details at the top-right. Cloning a pipeline allows you to create modifications without impacting the original pipeline. When cloning a pipeline, you become the owner of the cloned pipeline. When you clone a pipeline, you must give it a unique name because no duplicate names are allowed within all projects of the tenant. **So the** **name must be unique per tenant**. It is possible that you see the same pipeline name twice when a pipeline linked from another tenant is cloned with that same name in your tenant. The name is then still unique per tenant, but you will see them both in your tenant.

When you clone a Nextflow pipeline, a verification of the configured Nextflow version is done to prevent the use of deprecated versions.

### Documentation

The Documentation tab provides is the place where you explain how your pipeline works to users. The description appears in the tool repository but is excluded from exported CWL definitions. If no documentation has been provided, this tab will be empty.

### Definition (Graphical)

When using graphical mode for the pipeline definition, the Definition tab provides options for configuring the pipeline using a visualization panel and a list of component menus.

<table><thead><tr><th width="219">Menu</th><th>Description</th></tr></thead><tbody><tr><td>Machine profiles</td><td><a href="#compute">Compute types</a> available to use with Tools in the pipeline.</td></tr><tr><td>Shared settings</td><td>Settings for pipelines used in more than one tool.</td></tr><tr><td>Reference files</td><td>Descriptions of reference files used in the pipeline.</td></tr><tr><td>Input files</td><td>Descriptions of input files used in the pipeline.</td></tr><tr><td>Output files</td><td>Descriptions of output files used in the pipeline.</td></tr><tr><td>Tool</td><td>Details about the tool selected in the visualization panel.</td></tr><tr><td>Tool repository</td><td>A list of tools available to be used in the pipeline.</td></tr></tbody></table>

{% hint style="info" %}
In graphical mode, you can **drag and drop inputs** into the visualization panel to connect them to the tools. Make sure to **connect the input icons to the tool before editing the input details** in the component menu. Required tool inputs are indicated by a yellow connector.

**Safari is not supported** as browser for graphical editing.
{% endhint %}

{% hint style="warning" %}
When creating a graphical CWL pipeline, **do not use spaces in the input field names, use underscores instead**. The API performs normalization of input names when running the analysis to prevent issues with special characters (such as accented letters) by replacing them with their more common (unaccented) counterpart. Part of this normalization includes replacing spaces in names with underscores. This normalization is applied to file input name, reference file input name, step id and step parameters.

You will encounter the error ICA\_API\_004 "No value found for required input parameter" when trying to run an API analysis on a graphical pipeline that has been designed with spaces in input parameters.
{% endhint %}

### XML Configuration / JSON Inputform Files (Code)

This page is used to specify all relevant information about the pipeline parameters.

{% hint style="info" %}
There is a limit of 200 reports per report pattern which will be shown when you have multiple reports matching your regular expression.
{% endhint %}

### Compute Resources

#### Compute Nodes

For each process defined by the workflow, Platform Core will launch a compute node to execute the process.

* For each compute type, the `standard` (default - AWS on-demand) or `economy` (AWS spot instance) tiers can be selected.
* When selecting an **fpga** instance type for running analyses on Platform Core, it is recommended to use the medium size. While the large size offers slight performance benefits, these do not proportionately justify the associated cost increase for most use cases.
* When no type is specified, the default type of compute node is `standard-small`.

{% hint style="info" %}
You can see **which resources were used** in the different analysis steps at **Projects > your\_project > Flow > Analyses > your\_analysis > Steps tab**. (For child steps, these are displayed on the parent step)
{% endhint %}

By default, compute nodes have no scratch space. This is an advanced setting and should only be used when absolutely necessary as it will incur additional costs and may offer only limited performance benefits because it is not local to the compute node.

For simplicity and better integration, consider using shared storage available at `/ces`. It is what is provided in the Small/Medium/Large+ compute types. This shared storage is used when writing files with relative paths.

<details>

<summary>Scratch space notes</summary>

If you do require scratch space via a Nextflow pod annotation or a CWL resource requirement, the path is `/scratch`.

* For Nextflow `pod annotation: 'volumes.illumina.com/scratchSize', value: '1TiB'` will reserve 1 TiB.
* For CWL, adding `- class: ResourceRequirement tmpdirMin: 5000` to your requirements section will reserve 5000 MiB for CWL.

<mark style="color:red;">**Avoid the following**</mark> as it does not align with ICAv2 scratch space configuration.

* Container overlay tmp path: `/tmp`
* Legacy paths: `/ephemeral`
* Environment Variables ($TMPDIR, $TEMP and $TMP)
* Bash Command `mktemp`
* CWL `runtime.tmpdir`

</details>

#### Compute Types

Daemon sets and system processes consume approximately 1 CPU and 2 GB Memory from the base values shown in the table. Consumption will vary based on the activity of the pod.

| <p><br>Compute Type</p>      | CPUs | Mem (GiB) | Nextflow (`pod.value`) | CWL (`type, size`) |
| ---------------------------- | ---- | --------- | ---------------------- | ------------------ |
| standard-small               | 2    | 8         | standard-small         | standard, small    |
| standard-medium              | 4    | 16        | standard-medium        | standard, medium   |
| standard-large               | 8    | 32        | standard-large         | standard, large    |
| standard-xlarge              | 16   | 64        | standard-xlarge        | standard, xlarge   |
| standard-2xlarge             | 32   | 128       | standard-2xlarge       | standard, 2xlarge  |
| standard-3xlarge             | 64   | 256       | standard-3xlarge       | standard, 3xlarge  |
| hicpu-small                  | 16   | 32        | hicpu-small            | hicpu, small       |
| hicpu-medium                 | 36   | 72        | hicpu-medium           | hicpu, medium      |
| hicpu-large                  | 72   | 144       | hicpu-large            | hicpu, large       |
| himem-small                  | 8    | 64        | himem-small            | himem, small       |
| himem-medium                 | 16   | 128       | himem-medium           | himem, medium      |
| himem-large                  | 48   | 384       | himem-large            | himem, large       |
| himem-xlarge<sup>2</sup>     | 92   | 700       | himem-xlarge           | himem, xlarge      |
| hiio-small                   | 2    | 16        | hiio-small             | hiio, small        |
| hiio-medium                  | 4    | 32        | hiio-medium            | hiio, medium       |
| fpga2-medium<sup>1</sup>     | 24   | 256       | fpga2-medium           | fpga2,medium       |
| fpga2-large<sup>1</sup>      | 48   | 512       | fpga2-large            | fpga2,large        |
| gpu-small                    | 8    | 61        | gpu-small              | gpu, small         |
| gpu-medium                   | 32   | 244       | gpu-medium             | gpu, medium        |
| transfer-small<sup>3</sup>   | 4    | 10        | transfer-small         | transfer, small    |
| transfer-medium <sup>3</sup> | 8    | 15        | transfer-medium        | transfer, medium   |
| transfer-large<sup>3</sup>   | 16   | 30        | transfer-large         | transfer, large    |

{% hint style="warning" %} <sup>1</sup> **DRAGEN pipelines running on fpga2** compute type will incur a DRAGEN license cost of 0.10 BIC per gigabase of data processed, with additional **discounts** as shown below.

* **80 or less** gigabase per sample - no discount - 0.10 BIC per gigabase
* **> 80 to 160** gigabase per sample - 20% discount - 0.08 BIC per gigabase
* **> 160 to 240** gigabase per sample - 30% discount - 0.07 BIC per gigabase
* **> 240 to 320** gigabase per sample - 40% discount - 0.06 BIC per gigabase
* **> 320 and more** gigabase per sample - 50% discount - 0.05 BIC per gigabase

**DRAGEN Iterative gVCF Genotyper (iGG)** will incur a **license cost** of **0.6216 BIC per gigabase**. For example, a sample of 3.3 gigabase human reference will result in 2 BIC per sample. The associated **Compute costs** will be based on the compute instance chosen.

The **ORA (Original Read Archive) compression pipeline** is part of the DRAGEN platform. It performs lossless genomic data compression to reduce the size of FASTQ and FASTQ.GZ files (up to 4-6x smaller) while preserving data integrity with internal checksum verification. The ORA compression pipeline has a **license cost** of **0.017 BIC per input Gbase**; decompression does not have an associated license cost.
{% endhint %}

{% hint style="warning" %}
The ***DRAGEN\_Map\_Align*** **pipeline running on fpga2** has the standard **DRAGEN license cost** of 0.10 BIC per Gbase processed, with but replaces the standard volume discounts with the **discounts** shown below.

* **10 or less** gigabase per sample - no discount - 0.10 BIC per gigabase
* **> 10 to 25** gigabase per sample - 30% discount - 0.07 BIC per gigabase
* **> 25 to 60** gigabase per sample - 70% discount - 0.03 BIC per gigabase
* **> 60 and more** gigabase per sample - 85% discount - 0.015 BIC per gigabase
  {% endhint %}

{% hint style="info" %}
(2) The compute type **himem-xlarge** has low availability.
{% endhint %}

{% hint style="danger" %}
FPGA1 instances were decommissioned on Nov 1st 2025. Please migrate to F2 for improved capacity and performance with up to 40% reduced turnaround time for analysis.
{% endhint %}

{% hint style="info" %}
(3) The transfer size selected is based on the selected storage size for compute type and used during upload and download system tasks.
{% endhint %}

### Nextflow/CWL Files (Code)

Syntax highlighting is determined by the file type, but you can select alternative syntax highlighting with the drop-down selection list. The following formats are supported:

* DIFF (.diff)
* GROOVY (.groovy .nf)
* JAVASCRIPT (.js .javascript)
* JSON (.json)
* SH (.sh)
* SQL (.sql)
* TXT (.txt)
* XML (.xml)
* YAML (.yaml .cwl)

{% hint style="info" %}
If the file type is not recognized, it will default to text display. This can result in the application interpreting binary files as text when trying to display the contents.
{% endhint %}

#### Main.nf (Nextflow code)

The Nextflow project main script.

#### Nextflow\.config (Nextflow code)

The Nextflow configuration settings.

#### Workflow\.cwl (CWL code)

The Common Workflow Language main script.

#### Adding Files

Multiple files can be added by selecting the **+Create** option at the bottom of the screen to make pipelines more modular and manageable.

### Metadata Model

See [Metadata Models](/home/h-metadatamodels)

### Report

Here patterns for detecting report files in the analysis output can be defined. On opening an analysis result window of this pipeline, **an additional tab will display these report files.** The goal is to provide a pipeline-specific user-friendly representation of the analysis result.

To add a report select the **+ symbol** on the left side. Provide your report with a unique name, a regular expression matching the report and optionally, select the format of the report. This must be the source format of the report data generated during the analysis.

{% hint style="info" %}
There is a limit of 20 reports per report pattern which will be shown when you have multiple reports matching your regular expression.
{% endhint %}

***

## Start a New Analysis

Use the following instructions to start a new analysis for a single pipeline.

1. Select **Projects > your\_project > Flow > Pipelines.**
2. Select the pipeline or pipeline details of the pipeline you want to run.
3. Select **Start Analysis**.
4. Configure [analysis settings](#analysis-setting). (see below)
5. Select **Start Analysis**.
6. View the analysis status on the Analyses page.
   * **Requested**—The analysis is scheduled to begin.
   * **In Progress**—The analysis is in progress.
   * **Succeeded**—The analysis is complete.
   * **Failed** —The analysis has failed.
   * **Aborted** — The analysis was aborted before completing.
7. To end an analysis, select **Abort**.
8. To perform a completed analysis again, select **Re-run**.

#### Analysis Settings

The Start Analysis screen provides the configuration options for the analysis.

<table><thead><tr><th width="196">Field</th><th>Entry</th></tr></thead><tbody><tr><td>User Reference</td><td>The unique analysis name.</td></tr><tr><td>Pipeline</td><td>This is not editable, but provides a link to the pipeline so you want to look up details of the pipeline.</td></tr><tr><td>User tags (optional)</td><td>One or more tags used to filter the analysis list. Select from existing tags or type a new tag name in the field.</td></tr><tr><td>Notification (optional)</td><td>Enter your email address if you want to be notified when the analysis completes.</td></tr><tr><td>Output Folder<sup>1</sup></td><td>Select a folder in which the <strong>output folder of the analysis</strong> should be located. When <strong>no folder is selected</strong>, the output folder will be located in the <strong>root of the project</strong>.<br><br>When you open the folder selection dialog, you have the option to <strong>create a new folder</strong> (bottom of the screen). You can create nested folders by using the <code>folder/subfolder</code> syntax.<br><em>Do not use a / before the first folder or after the last subfolder in the folder creation dialog.</em></td></tr><tr><td>Logs Folder</td><td><p>Select a folder where the <strong>analysis logs</strong> will be stored. When <strong>no logs folder is selected</strong>, they will be stored as <strong>subfolder in the output folder</strong>. When you select a logs folder different from your output folder, the folders will be separated.<br><br>When you open the folder selection dialog, you can <strong>create a new folder</strong> (bottom of the screen). Create nested folders by using the <code>folder/subfolder</code> syntax. <em>Do not use a / before the first folder or after the last subfolder in the folder creation dialog.</em></p><p><em><strong>When choosing a folder containing data, files with the same name will be overwritten.</strong></em></p></td></tr><tr><td>Input</td><td>Select the input files to use in the analysis. (max. 50,000)</td></tr><tr><td>Settings (optional)</td><td>Provide input settings.</td></tr><tr><td>Resources</td><td>Select the storage size for your analysis. The available storage sizes depend on your selected Pricing subscription. See <a href="/pages/rewnkXVCIALBvAXMmNdn#data-storage">Storage</a> for more information.</td></tr></tbody></table>

<sup>1</sup> When using the API, you can [redirect analysis outputs](/project/p-flow/f-analyses#analysis-output-mappings) to be outside of the current project.

## Aborting Analyses

You can abort a running analysis from either the analysis overview (**Projects > your\_project > Flow > Analyses > your\_analysis > Manage > Abort**) or from the analysis details (**Projects > your\_project > Flow > Analyses > your\_analysis > Details tab > Abort**).

## View Analysis Results

You can view analysis results on the Analyses page or in the output folder on the Data page. You can also **rerun your analysis** from here.

1. Select a project, and then select the **Projects > your\_project > Flow > Analyses** page.
2. Select the desired analysis.
3. From the output files tab, expand the list if needed and select an output file.
   * If you want to **add or remove** any user or technical **tags**, you can do so from the data details view.
   * If you want to **download** the file, select Download.
   * To see the data in the data view where you can easily navigate between files, select **Open in data**.
4. To preview the file, select the **View** tab.

To see more details of your analyses, return to **Projects > your\_project > Flow > Analyses > your\_analysis**. The following tabs will be visible (depending on pipeline):

* **Details** - View information on the pipeline configuration.
* **Output files** - View the output of the Analysis.
* **Steps** - stderr and stdout information.
* **CWL** - The CWL pipeline definition.
* **Nextflow timeline** - Nextflow process execution timeline.
* **Nextflow execution** - Nextflow analysis report. Showing the run times, commands, resource usage and tasks for Nextflow analyses.
* **Report** - Shows the reports defined on the [pipeline report ](#report-code)tab.

{% hint style="info" %}
Other tabs can be available depending on your chosen pipeline.
{% endhint %}


# Nextflow

Platform Core supports running pipelines defined using [Nextflow](https://www.nextflow.io/). See [this tutorial](/tutorials/nextflow/nextflow-dragen-pipeline) for an example.

In order to run Nextflow pipelines, the following process-level attributes within the Nextflow definition must be considered.

## System Information

{% hint style="warning" %}
Pipelines using **Nextflow v20.10** will no longer run after April 22nd, 2026 and instead display "This pipeline is on an outdated nextflow version no longer supported on Platform Core". See the planned obsolescence notice below.
{% endhint %}

{% file src="/files/nr8QqPja4hWpQI3YlsZb" %}

<table data-header-hidden><thead><tr><th width="162.61328125"></th><th></th></tr></thead><tbody><tr><td>Nextflow version</td><td>22.04 (deprecated ⚠️), 24.10 (supported ✅), 25.10 (default ⭐)</td></tr><tr><td>Executor</td><td>Kubernetes</td></tr></tbody></table>

The following table shows when which Nextflow version is

* **Default** (⭐) This version will be proposed when creating a new Nextflow pipeline.
* **Supported** (✅) This version **can be selected** when you do not want the default Nextflow version.
* **Deprecated** (⚠️) This version can **not** be selected for **new pipelines**, but **pipelines** using this version will **still work**.
* **Removed** (❌). This version can not be selected when creating new pipelines and **pipelines** using this version will **no longer work**.

The switchover happens in the **January** release of that year.

<table><thead><tr><th width="116.43359375">Nextflow Version</th><th align="center">2026</th><th align="center">2027</th><th align="center">2028</th></tr></thead><tbody><tr><td>v20.10</td><td align="center">❌</td><td align="center">❌</td><td align="center">❌</td></tr><tr><td>v22.04</td><td align="center">⚠️</td><td align="center">⚠️</td><td align="center">❌</td></tr><tr><td>v24.10</td><td align="center">✅</td><td align="center">✅</td><td align="center">⚠️</td></tr><tr><td>v25.10</td><td align="center">⭐</td><td align="center">⭐</td><td align="center">✅</td></tr><tr><td>v26.10</td><td align="center">​</td><td align="center">✅</td><td align="center">⭐</td></tr><tr><td>v27.10</td><td align="center">​</td><td align="center">​</td><td align="center">✅</td></tr></tbody></table>

### Nextflow Version

You can select the Nextflow version while building a pipeline as follows:

<table data-header-hidden><thead><tr><th width="136.5">interface</th><th>Location</th></tr></thead><tbody><tr><td>GUI</td><td>Select the Nextflow version at <strong>Projects > your_project > flow > pipelines > your_pipeline > Details tab</strong>.</td></tr><tr><td>API</td><td>Select the Nextflow version by setting it in the optional field "<code>pipelineLanguageVersionId</code>". When not set, a default Nextflow version will be used for the pipeline.</td></tr></tbody></table>

## Compute Type

To specify a compute type for a Nextflow process, you can either define the cpu and memory (recommended) or use the compute type [predefined](/project/p-flow/f-pipelines#compute-types) sizes (required for specific hardware such as FPGA2).

{% hint style="info" %}
Do not mix these definition methods within the same process2, use either one or the other method.
{% endhint %}

### CPU and Memory

Specify the task resources using Nextflow directives in both the workflow script (.nf) and the configuration file (nextflow\.config) `cpus` defines the number of CPU cores allocated to the process, `memory` defines the amount of RAM which will be allocated.

**Process file** example

```nf
process ALIGN {
    cpus = 4
    memory = '16 GB'
    script:
    """
    your_command_here
    """
}
```

**Configuration file** example

```nextflow
process {
    withName: ALIGN {
        cpus = 4
        memory = '16 GB'
    }
}
```

Platform Core will convert the required resources to the correct predefined size. This enables porting public Nextflow pipelines without configuration changes.

### Predefined Sizes

To use the predefined sizes, use the [pod directive](https://www.nextflow.io/docs/latest/process.html#process-pod) within each process. Set the `annotation` to `scheduler.illumina.com/presetSize` and the `value` to the desired compute type. The default compute type, when this directive is not specified, is `standard-small` (2 CPUs and 8 GB of memory).

For example, if you want to use [FPGA 2 medium](/reference/r-pricing#compute), you need to add the line below

```groovy
pod annotation: 'scheduler.illumina.com/presetSize', value: 'fpga2-medium'
```

{% hint style="info" %}
Often, there is a need to select the compute size for a process dynamically based on user input and other factors. The Kubernetes executor used on Platform Core does not use the `cpu` and `memory`directives, so instead, you can dynamically set the `pod` directive, as mentioned [here](https://www.nextflow.io/docs/latest/process.html#dynamic-directives). e.g.

```groovy
process foo {
    // Assuming that params.compute_size is set to a valid size such as 'standard-small', 'standard-medium', etc.
    pod annotation: 'scheduler.illumina.com/presetSize', value: "${params.compute_size}"
}
```

It can also be specified in the [configuration file](https://www.nextflow.io/docs/latest/config.html). See the example configuration below:

```groovy
// Set the default pod
pod = [
    annotation: 'scheduler.illumina.com/presetSize',
    value     : 'standard-small'
]

withName: 'big_memory_process' {
    pod = [
        annotation: 'scheduler.illumina.com/presetSize',
        value     : 'himem-large'
    ]
}

// Use an FPGA2 instance for dragen processes
withLabel: 'dragen' {
    pod = [
        annotation: 'scheduler.illumina.com/presetSize',
        value     : 'fpga2-medium'
    ]
}
```

{% endhint %}

### Standard vs Economy

#### Concept

For each compute type, you can choose between the

* `scheduler.illumina.com/lifecycle: standard` - [**AWS on-demand**](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-on-demand-instances.html) (Default) or
* `scheduler.illumina.com/lifecycle: economy` - [**AWS spot instance**](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-spot-instances.html) tiers.

<table><thead><tr><th width="124.84375"></th><th>On-Demand Instance</th><th>Spot Instance</th></tr></thead><tbody><tr><td>Pricing</td><td>Fixed <a href="/pages/rewnkXVCIALBvAXMmNdn">price</a> per second with 60-second minimum.</td><td><a href="/pages/rewnkXVCIALBvAXMmNdn">Cheaper</a> than On-Demand.</td></tr><tr><td>Availability</td><td>Guaranteed capacity with Full control of starting, stopping, and terminating.</td><td>Not guaranteed. Depends on unused AWS capacity. Can be terminated and reclaimed by AWS when the capacity is needed for other processes with <strong>2 minutes notice</strong>.</td></tr><tr><td>Best for</td><td>Ideal for critical workloads and urgent scaling needs.</td><td>Best for cost optimization and non-critical workloads as interruptions can occur any time.</td></tr></tbody></table>

#### Configuration

You can switch to economy in the process itself with the pod directive or in the nextflow\.config file.

Process example

```groovy
process foo {
    pod annotation: 'scheduler.illumina.com/lifecycle', value: "economy"
}
```

nextlow\.config example

```groovy
process.withName: PROCESS_NAME {
    pod.annotations = [
        'scheduler.illumina.com/lifecycle': 'economy'
    ]
}
```

## Inputs

Inputs are specified via the [JSON-based input](/project/p-flow/f-pipelines/json-based-input-forms) form or [XML input form](/project/p-flow/f-pipelines/pi-inputform). The specified `code` in the XML will correspond to the field in the `params` object that is available in the workflow. Refer to the [tutorial](/tutorials/nextflow/nextflow-dragen-pipeline) for an example.

## Outputs

Outputs for Nextflow pipelines are uploaded from the `out` folder in the attached shared filesystem.

The [`publishDir` directive](https://www.nextflow.io/docs/latest/process.html#publishdir) can be used to **symlink** (recommended for publishing files and folders to single destinations), **link** (recommended for publishing the same file to multiple locations) copy or move data to the correct folder.

* **Symlinking** is fast and does not increase storage cost as it creates a file **pointer** instead of copying or moving data. It can handle both individual **files and folders**, but it can **not publish the same data to multiple destinations** and is **not resistant to deletion of the original file.**
* **Linking** creates a hard link to the file and so does not increase storage cost either. It also keeps working after deletion (removing the reference) of the original file. Linking **supports** publishing the **same file to multiple destinations.** It can be used for files, but **not folders** and the **files must be on the same filesystem.**

Data will be uploaded to the Platform Core project after the pipeline execution completes.

```groovy
publishDir 'out', mode: 'symlink'
```

```groovy
publishDir 'out', mode: 'link'
```

<details>

<summary>Nextflow version 20.10.10 (Deprecated)</summary>

**Version 20.10 will be obsoleted on April 22nd, 2026. After this date, all existing pipelines using Nextflow v20.10 will no longer be able to run.**

For Nextflow version 20.10.10 on Platform Core, using the "copy" method in the `publishDir` directive for uploading output files that consume large amounts of storage may cause workflow runs to complete with missing files. The underlying issue is that file uploads may silently fail (without any error messages) during the `publishDir` process due to insufficient disk space, resulting in incomplete output delivery.

Solutions:

1. Use "[symlink](https://help.ica.illumina.com/project/p-flow/f-pipelines/pi-nextflow#outputs)" instead of "copy" in the `publishDir` directive. Symlinking creates a link to the original file rather than copying it, which doesn’t consume additional disk space. This can prevent the issue of silent file upload failures due to disk space limitations.
2. Use Nextflow 22.04 or later and enable the "[failOnError](https://www.nextflow.io/docs/latest/process.html#publishdir)" `publishDir` option. This option ensures that the workflow will fail and provide an error message if there's an issue with publishing files, rather than completing silently without all expected outputs.

</details>

## Nextflow Configuration

During execution, the Nextflow pipeline runner determines the environment settings based on values passed via the command-line or via a configuration file (see [Nextflow Configuration documentation](https://www.nextflow.io/docs/latest/config.html)). When creating a Nextflow pipeline, use the nextflow\.config tab in the UI (or API) to specify a nextflow configuration file to be used when launching the pipeline.

Syntax highlighting is determined by the file type, but you can select alternative syntax highlighting with the drop-down selection list.

![nextflowconfig-0](/files/pB5gUt5oyOWwGRnLd5d3)

{% hint style="info" %}
If no Docker image is specified, Ubuntu will be used as default.
{% endhint %}

The following configuration settings will be ignored if provided as they are overridden by the system:

```yaml
executor.name
executor.queueSize
k8s.namespace
k8s.serviceAccount
k8s.launchDir
k8s.projectDir
k8s.workDir
k8s.storageClaimName
k8s.storageMountPath
trace.enabled
trace.file
trace.fields
timeline.enabled
timeline.file
report.enabled
report.file
dag.enabled
dag.file
```

## Best Practices

### Process Time

Setting a timeout to between 2 and 4 times the expected processing time with the [**time**](https://www.nextflow.io/docs/latest/reference/process.html#process-time) directive for processes or task will ensure that no stuck processes remain indefinitely. Stuck process keep incurring costs for the occupied resources, so if the process can not complete within that timespan, it is safer and more economical to end the process and retry.

### Sample Sheet File Ingestion

When you want to use a sample sheet with references to files as Nextflow input, add an extra input to the pipeline. This extra input lets the user select the samplesheet-mentioned files from their project. At run time, those files will get staged in the working directory, and when Nextflow parses the samplesheet and looks for those files without paths, it will find them there. You can not use file paths in a sample sheet without selecting the files in the input form because files are only passed as file/folder ids in the API payload when the analysis is launched.

You can include public data such as **http urls** because Nextflow is able download those. Nextflow is also able to download publicly accessible **S3 urls** (s3://...). You can **not** use Illumina's **urn**:ilmn:ica:region:... structure.

## Migration

### Migrating from FPGA to FPGA2

New versions of existing **DRAGEN workflows** have been created to **support F2 (FPGA2)** instances as F1 (FPGA) instances have been decommissioned. Please consult the [DRAGEN BSSH/ICA end of life roadmap](https://help.dragen.illumina.com/reference/eol-transition#dragen-bssh-ica-workflows-end-of-life-roadmap) for more information. You will need to migrate your pipelines from FPGA to FPGA2.

As long as your pipeline is still in [draft](/project/p-flow/f-pipelines#pipeline-statuses) status, you can update them with the FPGA2 configuration, but once the pipeline has been [released](/project/p-flow/f-pipelines#pipeline-statuses), you need to **clone and edit the pipeline** as it is protected against editing. Cloning the pipeline is done at **projects > your\_project > Flow > Pipelines > open pipeline details > Clone** (top right). Setting the compute resources can be done in the **.nf** file directly or in the **nextflow\.config** file.

#### .nf

```
process DRAGEN_PROCESS {
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'fpga2-medium'
 
    script:
    """
    your_command_here
    """
}
```

#### nextflow\.config

```
process {
    withLabel: 'dragen' {
        pod = [
            annotation: 'scheduler.illumina.com/presetSize',
            value     : 'fpga2-medium'
        ]
    }
}
```

### Entrypoint

Starting with Nextflow version 22.08, Nextflow no longer overrides the container ENTRYPOINT with `/bin/bash`. Previously, Nextflow would force `bin/bash` as the entrypoint regardless of what the image defined, which meant environment setup in custom ENTRYPOINTs (such as conda activation scripts) was silently being bypassed. As of Nextflow 22.08, the container's native ENTRYPOINT is respected.

Containers relying on a custom ENTRYPOINT to set up the environment (e.g., activating a conda environment or prepending to PATH) will no longer have `/bin/bash` injected as the shell. If tools are not directly callable from the default shell PATH, you will encounter "command not found" errors.

To verify that a container image is compatible, run the following command locally for each tool used in your pipeline:

{% code overflow="wrap" %}

```bash
docker run <image> <tool> --help
```

{% endcode %}

This command must complete successfully (exit code 0 and display the tool help text). If it fails with "command not found", the image needs to be updated so that the tool is directly callable without relying on ENTRYPOINT for environment setup.

#### Set PATH directly in the Dockerfile (Recommended)

{% code overflow="wrap" %}

```docker
ENV PATH="/opt/conda/envs/myenv/bin:$PATH"  
```

{% endcode %}

#### Use absolute paths in your pipeline

This solution can be used for quick fixes, but has a higher maintenance cost as it uses hardcoded paths.

{% code overflow="wrap" %}

```nextflow
process MY_PROCESS {
    script:
       /opt/conda/envs/myenv/bin/bcftools view input.vcf > output.vcf
}
```

{% endcode %}

#### Activate the environment in beforeScript

This mimics the legacy behavior of what ENTRYPOINT used to do.

{% code overflow="wrap" %}

```nextflow
process {
    beforeScript = 'source /opt/conda/etc/profile.d/conda.sh && conda activate myenv'
}
```

{% endcode %}

#### Restore legacy behavior (deprecated — not recommended)

The old behavior can be temporarily restored by setting the environment variable

{% code overflow="wrap" %}

```
NXF_CONTAINER_ENTRYPOINT_OVERRIDE=true
```

{% endcode %}

**This option is** [**deprecated**](https://docs.seqera.io/nextflow/reference/env-vars#nxf_container_entrypoint_override) and can be removed without notice in any future Nextflow release, at which point any pipeline depending on it will no longer function.


# CWL

Platform Core supports running pipelines defined using [Common Workflow Language (CWL)](https://www.commonwl.org/).

## Compute Type

To specify a compute type for a CWL CommandLineTool, either define the **ram and number of cores** **or** use the **resource type and size**. The Platform Core compute type will automatically be determined based on [CWL ResourceRequirement](https://www.commonwl.org/v1.0/CommandLineTool.html#ResourceRequirement) coresMin/coresMax (CPU) and ramMin/ramMax (Memory) values using a "best fit" strategy to meet the minimum specified requirements (See the [Compute Types](/project/p-flow/f-pipelines#compute-types) table to see to what the resources are mapped).

For example, take the following `ResourceRequirements`:

```yaml
requirements:
    ResourceRequirement:
      ramMin: 10240
      coresMin: 6
```

This will result in a best fit of `standard-large` Platform Core compute type request for the task.

{% hint style="info" %}
If the specified requirements can not be met by any of the presets, the task will be rejected and failed.
{% endhint %}

See the example below to use the `ResourceRequirement` in the cwl workflow with the Predefined [Compute Types](/project/p-flow/f-pipelines#compute-types)

```yaml
requirements:
    ResourceRequirement:
        https://platform.illumina.com/rdf/ica/resources:type: fpga2
        https://platform.illumina.com/rdf/ica/resources:size: medium 
        https://platform.illumina.com/rdf/ica/resources:tier: standard
```

{% hint style="info" %}

* FPGA requirements can not be set by means of CWL ResourceRequirements.
* The Machine Profile Resource in the graphical editor will override whatever is set for requirements in the ResourceRequirement.
  {% endhint %}

### Standard vs Economy

For each compute type, you can choose between the

* Standard - [**AWS on-demand**](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-on-demand-instances.html) (Default) or
* Economy - [**AWS spot instance**](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-spot-instances.html) tiers.

You can set economy mode with the "tier" parameter

```yaml
requirements:
    ResourceRequirement:
        https://platform.illumina.com/rdf/ica/resources:type: himem
        https://platform.illumina.com/rdf/ica/resources:size: small 
        https://platform.illumina.com/rdf/ica/resources:tier: economy
```

## Considerations

> If no Docker image is specified, Ubuntu will be used as default. Both : and / can be used as separator.

## CWL Overrides

Platform Core supports overriding workflow requirements at load time using Command Line Interface (CLI) with JSON input. Please refer to [CWL documentation](https://github.com/common-workflow-language/cwltool/tree/3.0.20201203173111#overriding-workflow-requirements-at-load-time) for more details on the CWL overrides feature.

In Platform Core you can provide the "override" recipes as a part of the input JSON. The following example uses CWL overrides to change the environment variable requirement at load time.

```bash
icav2 projectpipelines start cwl cli-tutorial --data-id fil.a725a68301ee4e6ad28908da12510c25 --input-json '{
  "ipFQ": {
    "class": "File",
    "path": "test.fastq"
  },
  "cwltool:overrides": {
  "tool-fqTOfa.cwl": {
    "requirements": {
      "EnvVarRequirement": {
        "envDef": {
          "MESSAGE": "override_value"
          }
        }                                       
       }
      }
    }
}' --type-input JSON --user-reference overrides-example
```


# XML Input Form

Pipelines defined using the "Code" mode require either an XML-based or JSON-based input form to define the fields shown on the launch view in the user interface (UI). The XML-based input form is defined in the "XML Configuration" tab of the pipeline editing view.

<figure><img src="/files/gZKFQC35Sjcy9f5Pgl1x" alt=""><figcaption></figcaption></figure>

The input form XML must adhere to the input form schema.

## Empty Form

During the creation of a Nextflow pipeline the user is given an empty form to fill out.

{% code overflow="wrap" %}

```xml
<pipeline code="" version="1.0" xmlns="xsd://www.illumina.com/ica/cp/pipelinedefinition">
    <dataInputs>
    </dataInputs>
    <steps>
    </steps>
</pipeline>
```

{% endcode %}

## Files

The input files are specified within a single **DataInputs** node. An individual input is then specified in a separate **DataInput** node. A **DataInput** node contains following attributes:

* *code*: an unique id. Required.
* *format*: specifying the format of the input: FASTA, TXT, JSON, UNKNOWN, etc. Multiple entries are possible: example below. Required.
* *type*: is it a FILE or a DIRECTORY? Multiple entries are not allowed. Required.
* *required*: is this input required for the execution of a pipeline? Required.
* *multiValue*: are multiple files as an input allowed? Required.
* *dataFilter*: TBD. Optional.

Additionally, **DataInput** has two elements: *label* for labelling the input and *description* for a free text description of the input.

### Single file input

An example of a single file input which can be in a TXT, CSV, or FASTA format.

{% code overflow="wrap" %}

```xml
        <pd:dataInput code="in" format="TXT, CSV, FASTA" type="FILE" required="true" multiValue="false">
            <pd:label>Input file</pd:label>
            <pd:description>Input file can be either in TXT, CSV or FASTA format.</pd:description>
        </pd:dataInput>
```

{% endcode %}

### Folder as an input

To use a folder as an input the following form is required:

{% code overflow="wrap" %}

```xml
    <pd:dataInput code="fastq_folder" format="UNKNOWN" type="DIRECTORY" required="false" multiValue="false">
         <pd:label>fastq folder path</pd:label>
        <pd:description>Providing Fastq folder</pd:description>
    </pd:dataInput>
```

{% endcode %}

### Multiple files as an input

For multiple files, set the attribute *multiValue* to true. This will make it so the variable is considered to be of **type list \[]**, so adapt your pipeline when changing from single value to multiValue.

{% code overflow="wrap" %}

```xml
<pd:dataInput code="tumor_fastqs" format="FASTQ" type="FILE" required="false" multiValue="true">
    <pd:label>Tumor FASTQs</pd:label>
    <pd:description>Tumor FASTQ files to be provided as input. FASTQ files must have "_LXXX" in its filename to denote the lane and "_RX" to denote the read number. If either is omitted, lane 1 and read 1 will be used in the FASTQ list. The tool will automatically write a FASTQ list from all files provided and process each sample in batch in tumor-only mode. However, for tumor-normal mode, only one sample each can be provided.
    </pd:description>
</pd:dataInput>
```

{% endcode %}

## Settings

Settings (as opposed to files) are specified within the **steps** node. Settings represent any non-file input to the workflow, including but not limited to, strings, booleans, integers, etc. The following hierarchy of nodes must be followed: *steps* > *step* > *tool* > *parameter*. The *parameter* node must contain following attributes:

* *code*: unique id. This is the parameter name that is passed to the workflow
* *minValues*: how many values (at least) should be specified for this setting. If this setting is required, `minValues` should be set to 1.
* *maxValues*: how many values (at most) should be specified for this setting
* *classification*: is this setting specified by the user?

In the code below a string setting with the identifier *inp1* is specified.

{% code overflow="wrap" %}

```xml
    <pd:steps>
        <pd:step execution="MANDATORY" code="General">
            <pd:label>General</pd:label>
            <pd:description>General parameters</pd:description>
            <pd:tool code="generalparameters">
                <pd:label>generalparameters</pd:label>
                <pd:description></pd:description>
                <pd:parameter code="inp1" minValues="1" maxValues="3" classification="USER">
                    <pd:label>inp1</pd:label>
                    <pd:description>first</pd:description>
                    <pd:stringType/>
                    <pd:value></pd:value>
                </pd:parameter>
            </pd:tool>
        </pd:step>
    </pd:steps>
```

{% endcode %}

Examples of the following types of settings are shown in the subsequent sections. Within each type, the `value` tag can be used to denote a default value in the UI, or can be left blank to have no default. Note that setting a default value has **no impact on analyses launched via the API.**

### Integers

For an integer setting the following schema with an element *integerType* is to be used. To define an allowed range use the attributes *minimumValue* and *maximumValue*.

{% code overflow="wrap" %}

```xml
<pd:parameter code="ht_seed_len" minValues="0" maxValues="1" classification="USER">
    <pd:label>Seed Length</pd:label>
    <pd:description>Initial length in nucleotides of seeds from the reference genome to populate into the hash table. Consult the DRAGEN manual for recommended lengths. Corresponds to DRAGEN argument --ht-seed-len.
    </pd:description>
    <pd:integerType minimumValue="10" maximumValue="50"/>
    <pd:value>21</pd:value>
</pd:parameter>
```

{% endcode %}

### Options

Options types can be used to designate options from a drop-down list in the UI. The selected option will be passed to the workflow as a string. This currently has no impact when launching from the API, however.

{% code overflow="wrap" %}

```xml
<pd:parameter code="cnv_segmentation_mode" minValues="0" maxValues="1" classification="USER">
    <pd:label>Segmentation Algorithm</pd:label>
    <pd:description> DRAGEN implements multiple segmentation algorithms, including the following algorithms, Circular Binary Segmentation (CBS) and Shifting Level Models (SLM).
    </pd:description>
    <pd:optionsType>
        <pd:option>CBS</pd:option>
        <pd:option>SLM</pd:option>
        <pd:option>HSLM</pd:option>
        <pd:option>ASLM</pd:option>
    </pd:optionsType>
    <pd:value>false</pd:value>
</pd:parameter>
```

{% endcode %}

Option types can also be used to specify a boolean, for example

{% code overflow="wrap" %}

```xml
<pd:parameter code="output_format" minValues="1" maxValues="1" classification="USER">
    <pd:label>Map/Align Output</pd:label>
    <pd:description></pd:description>
    <pd:optionsType>
        <pd:option>BAM</pd:option>
        <pd:option>CRAM</pd:option>
    </pd:optionsType>
    <pd:value>BAM</pd:value>
</pd:parameter>
```

{% endcode %}

### Strings

For a string setting the following schema with an element `stringType` is to be used.

{% code overflow="wrap" %}

```xml
<pd:parameter code="output_file_prefix" minValues="1" maxValues="1" classification="USER">
    <pd:label>Output File Prefix</pd:label>
    <pd:description></pd:description>
    <pd:stringType/>
    <pd:value>tumor</pd:value>
</pd:parameter>
```

{% endcode %}

### Booleans

For a boolean setting, `booleanType` can be used.

```xml
<pd:parameter code="quick_qc" minValues="0" maxValues="1" classification="USER">
    <pd:label>quick_qc</pd:label>
    <pd:description></pd:description>
    <pd:booleanType/>
    <pd:value></pd:value>
</pd:parameter>
```

## Limitations

One known limitation of the schema presented above is the inability to specify a parameter that can be multiple type, e.g. File or String. One way to implement this requirement would be to define two optional parameters: one for File input and the second for String input. At the moment Platform Core UI doesn't validate whether at least one of these parameters is populated - this check can be done within the pipeline itself.

Below one can find both a main.nf and XML configuration of a generic pipeline with two optional inputs. One can use it as a template to address similar issues. If the *file* parameter is set, it will be used. If the *str* parameter is set but *file* is not, the *str* parameter will be used. If neither of both is used, the pipeline aborts with an informative error message.

```groovy
nextflow.enable.dsl = 2

// Define parameters with default values
params.file = false
params.str = false

// Check that at least one of the parameters is specified
if (!params.file && !params.str) {
    error "You must specify at least one input: --file or --str"
}

process printInputs {
    
    container 'public.ecr.aws/lts/ubuntu:22.04'
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'

    input:
    file(input_file)

    script:
    """
    echo "File contents:"
    cat $input_file
    """
}

process printInputs2 {

    container 'public.ecr.aws/lts/ubuntu:22.04'
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'

    input:
    val(input_str)

    script:
    """
    echo "String input: $input_str"
    """
}

workflow {
    if (params.file) {
        file_ch = Channel.fromPath(params.file)
        file_ch.view()
        str_ch = Channel.empty()
        printInputs(file_ch)
    }
    else {
        file_ch = Channel.empty()
        str_ch = Channel.of(params.str)
        str_ch.view()
        file_ch.view()
        printInputs2(str_ch)
    } 
}
```

```xml
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition" code="" version="1.0">
    <pd:dataInputs>
        <pd:dataInput code="file" format="TXT" type="FILE" required="false" multiValue="false">
            <pd:label>in</pd:label>
            <pd:description>Generic file input</pd:description>
        </pd:dataInput>
    </pd:dataInputs>
    <pd:steps>
        <pd:step execution="MANDATORY" code="general">
            <pd:label>General Options</pd:label>
            <pd:description locked="false"></pd:description>
            <pd:tool code="general">
                <pd:label locked="false"></pd:label>
                <pd:description locked="false"></pd:description>
                <pd:parameter code="str" minValues="0" maxValues="1" classification="USER">
                    <pd:label>String</pd:label>
                    <pd:description></pd:description>
                    <pd:stringType/>
                    <pd:value>string</pd:value>
                </pd:parameter>
            </pd:tool>
        </pd:step>
    </pd:steps>
</pd:pipeline>
```


# JSON-Based input forms

### Introduction

Pipelines defined using the "Code" mode require an XML or JSON-based input form to define the fields shown on the launch view in the user interface (UI).

To create a JSON-based Nextflow (or CWL) pipeline, go to **Projects > your\_project > Flow > Pipelines > +Create > Nextflow (or CWL) > JSON-based**.

The files listed here, located on the inputform files tab, work together for evaluating and presenting JSON-based input.

* [**inputForm.json**](/project/p-flow/f-pipelines/json-based-input-forms/inputform-json) contains the actual input form which is rendered when starting the pipeline run.
* [**onRender.js**](/project/p-flow/f-pipelines/json-based-input-forms/onrender.js) is triggered when a value is changed.
* [**onSubmit.js**](/project/p-flow/f-pipelines/json-based-input-forms/onsubmit.js) is triggered when starting a pipeline via the GUI or API.
* [**onFinish.js**](/project/p-flow/f-pipelines/json-based-input-forms/onfinished.js) (optional, not automatically created) is triggered once the pipeline analysis has been completed

Use **+ Create** to add additional files and **Simulate** to test your inputForms.

Scripting execution supports crossfield validation of the values, hiding fields, making them required, .... based on value changes.


# inputForm.json

### inputForm.json

The JSON schema allowing you to define the input parameters. See the [inputForm.json](/project/p-flow/f-pipelines/json-based-input-forms/inputform-json/inputform.json-syntax) page for syntax details.

{% hint style="info" %}
The inputForm.json file has a size limit of 10 MB and a maximum of 200 fields.
{% endhint %}

## Parameter types <a href="#pipelineinputforms-parametertypes" id="pipelineinputforms-parametertypes"></a>

<table><thead><tr><th width="181">Type</th><th>Usage</th></tr></thead><tbody><tr><td>textbox</td><td>Corresponds to stringType in xml.</td></tr><tr><td>checkbox</td><td>A checkbox that supports the option of being required, so can serve as an active consent feature. (corresponds to the booleanType in xml).</td></tr><tr><td>radio</td><td>A radio button group to select one from a list of choices. The values to choose from must be unique.</td></tr><tr><td>select</td><td>A dropdown selection to select one from a list of choices. This can be used for both single-level lists and tree-based lists.</td></tr><tr><td>number</td><td>The value is of Number type in javascript and Double type in java. (corresponds to doubleType in xml).</td></tr><tr><td>integer</td><td>Corresponds to java Integer.</td></tr><tr><td>data</td><td>Data such as files.</td></tr><tr><td>section</td><td>For splitting up fields, to give structure. Rendered as subtitles. No values are to be assigned to these fields.</td></tr><tr><td>text</td><td>To display informational messages. No values are to be assigned to these fields.</td></tr><tr><td>fieldgroup</td><td>Can contain parameters or other groups. Allows to have repeating sets of parameters, for instance when a father|mother|child choice needs to be linked to each file input. So if you want to have the same elements multiple times in your form, combine them into a fieldgroup.<br>Does not support the emptyValuesAllowed attribute.</td></tr></tbody></table>

## Parameter Attributes <a href="#pipelineinputforms-parameterattributes" id="pipelineinputforms-parameterattributes"></a>

These attributes can be used to configure all parameter types.

<table><thead><tr><th width="251">Attribute</th><th>Purpose</th></tr></thead><tbody><tr><td>label</td><td>The display label for this parameter. Optional but recommended, id will be used if missing.</td></tr><tr><td>minValues</td><td>The minimal amount of values that needs to be present. Default when not set is 0. Set to >=1 to make the field required.</td></tr><tr><td>maxValues</td><td>The maximal amount of values that need to be present. Default when not set is 1.</td></tr><tr><td>minMaxValuesMessage</td><td>The error message displayed when minValues or maxValues is not adhered to. When not set, a default message is generated.</td></tr><tr><td>helpText</td><td>A helper text about the parameter. Will be displayed in smaller font with the parameter.</td></tr><tr><td>placeHolderText</td><td>An optional short hint ( a word or short phrase) to aid the user when the field has no value.</td></tr><tr><td>value</td><td>The value of the parameter. Can be considered default value.</td></tr><tr><td>minLength</td><td>Only applied on type="textbox". Value is a positive integer.</td></tr><tr><td>maxLength</td><td>Only applied on type="textbox". Value is a positive integer.</td></tr><tr><td>min</td><td><p>Minimal allowed value for '<strong>integer</strong>' and '<strong>number</strong>' type.</p><ul><li>for 'integer' type fields the minimal and maximal values are -100000000000000000 and 100000000000000000.</li><li>for 'number' type fields the max precision is 15 significant digits and the exponent needs to be between -300 and +300.</li></ul></td></tr><tr><td>max</td><td><p>Maximal allowed value for '<strong>integer</strong>' and '<strong>number</strong>' type.</p><ul><li>for 'integer' type fields the minimal and maximal values are -100000000000000000 and 100000000000000000.</li><li>for 'number' type fields the max precision is 15 significant digits and the exponent needs to be between -300 and +300.</li></ul></td></tr><tr><td>choices</td><td>A list of choices with for each a "value", "text" (is label), "selected" (only 1 true supported), "disabled". "parent" can be used to build hierarchical choicetrees. "availableWhen" can be used for conditional presence of the choice based on values of other fields. Parent and value must be unique, you can not use the same value for both.</td></tr><tr><td>fields</td><td>The list of sub fields for type fieldgroup.</td></tr><tr><td>dataFilter</td><td>For defining the filtering when type is 'data'. Use <strong>nameFilter</strong> for matching the name of the file, <strong>dataFormat</strong> for file format and <strong>dataType</strong> for selecting between files and directories. (<em>To see the data formats, open the file details in Platform Core and look at the Format on the data details. You can expand the dropdown list to see the syntax.)</em> The <strong>dataType="file"</strong> also accepts S3 and HTTP(S) URLs.</td></tr><tr><td>regex</td><td>The regex pattern the value must adhere to. Only applied on type="textbox".</td></tr><tr><td>regexErrorMessage</td><td>The optional error message when the value does not adhere to the "regex". A default message will be used if this parameter is not present. It is highly recommended to set this as the default message will show the regex which is typically very technical.</td></tr><tr><td>hidden</td><td>Makes this parameter hidden. Can be made visible later in onRender.js or can be used to set hardcoded values of which the user should be aware.</td></tr><tr><td>disabled</td><td>Shows the parameter but makes editing it impossible. The value can still be altered by onRender.js.</td></tr><tr><td>emptyValuesAllowed</td><td>When maxValues is 1 or not set and emptyValuesAllowed is true, the values may contain null entries. Default is false.</td></tr><tr><td>updateRenderOnChange</td><td>When true, the onRender javascript function is triggered each time the user changes the value of this field. Default is false.</td></tr><tr><td>dropValueWhenDisabled</td><td>When this is present and true and the field has <em>disabled</em> being true, then the value will be omitted during the submit handling (on the onSubmit result).</td></tr></tbody></table>

<details>

<summary>Tree structure example</summary>

"choices" can be used for a single list or for a tree-structured list. See below for an example for how to set up a tree structure.

```json
{
  "fields": [
    {
      "id": "myTreeList",
      "type": "select",
      "label": "Selection Tree Example",
      "choices": [
        {
          "text": "trunk",
          "value": "treetrunk"
        },
        {
          "text": "branch",
          "value": "treebranch",
          "parent":"treetrunk"
        },
        {
          "text": "leaf",
          "value": "treeleaf",
          "parent":"treebranch"
        },
        {
          "text": "bird",
          "value": "happybird",
          "parent":"treebranch"
        },
        {
          "text": "cat",
          "value": "cat",
          "parent": "treetrunk",
          "disabled": true
        }
      ],
      "minValues": 1,
      "maxValues": 3,
      "helpText": "This is a tree example"
    }
  ]
}
```

</details>

## Experimental Features

<table><thead><tr><th width="267">Feature</th><th></th></tr></thead><tbody><tr><td>Streamable inputs</td><td>Adding <code>"streamable":true</code> to an input field of type "<strong>data</strong>" makes it a streamable input.</td></tr></tbody></table>


# JSON Schema

In the InputForm.json, use the syntax for the individual components you want to as listed below. This is a listing of all the currently available components and not to be used "as is", but adapted to the inputs you need in your inputform. For more information on JSON schema syntax, please see the [json-schema website](https://json-schema.org/).

{% code fullWidth="true" %}

```json
{
  "$id": "#ica-pipeline-input-form",
  "$schema": "http://json-schema.org/draft-07/schema#",
  "title": "ICA Pipeline Input Forms",
  "description": "Describes the syntax for defining input setting forms for ICA pipelines",
  "type": "object",
  "additionalProperties": false,
  "properties": {
    "fields": {
      "description": "The list of setting fields",
      "type": "array",
      "items": {
        "$ref": "#/definitions/ica_pipeline_input_form_field"
      }
    }
  },
  "required": [
    "fields"
  ],
  "definitions": {
    "ica_pipeline_input_form_field": {
      "$id": "#ica_pipeline_input_form_field",
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "id": {
          "description": "The unique identifier for this field. Will be available with this key to the pipeline script.",
          "type": "string",
          "pattern": "^[a-zA-Z-0-9\\-_\\.\\s\\+\\[\\]]+$"
        },
        "type": {
          "type": "string",
          "enum": [
            "textbox",
            "checkbox",
            "radio",
            "select",
            "number",
            "integer",
            "data",
            "section",
            "text",
            "fieldgroup"
          ]
        },
        "label": {
          "type": "string"
        },
        "minValues": {
          "description": "The minimal amount of values that needs to be present. Default is 0 when not provided. Set to >=1 to make the field required.",
          "type": "integer",
          "minimum": 0
        },
        "maxValues": {
          "description": "The maximal amount of values that needs to be present. Default is 1 when not provided.",
          "type": "integer",
          "exclusiveMinimum": 0
        },
        "minMaxValuesMessage": {
          "description": "The error message displayed when minValues or maxValues is not adhered to. When not provided a default message is generated.",
          "type": "string"
        },
        "helpText": {
          "type": "string"
        },
        "placeHolderText": {
          "description": "An optional short hint (a word or short phrase) to aid the user when the field has no value.",
          "type": "string"
        },
        "value": {
         "description": "The value for the field. Can be an array for multi-value fields. For 'number' type values the exponent needs to be between -300 and +300 and max precision is 15. For 'integer' type values the value needs to between -100000000000000000 and 100000000000000000"
         },
        "minLength": {
          "type": "integer",
          "minimum": 0
        },
        "maxLength": {
          "type": "integer",
          "exclusiveMinimum": 0
        },
        "min": {
          "description": "Minimal allowed value for 'integer' and 'number' type. Exponent needs to be between -300 and +300 and max precision is 15.",
          "type": "number"
        },
        "max": {
          "description": "Maximal allowed value for 'integer' and 'number' type. Exponent needs to be between -300 and +300 and max precision is 15.",
          "type": "number"
        },
        "choices": {
          "type": "array",
          "items": {
            "$ref": "#/definitions/ica_pipeline_input_form_field_choice"
          }
        },
        "fields": {
          "description": "The list of setting sub fields for type fieldgroup",
          "type": "array",
          "items": {
            "$ref": "#/definitions/ica_pipeline_input_form_field"
          }
        },
        "dataFilter": {
          "description": "For defining the filtering when type is 'data'.",
          "type": "object",
          "additionalProperties": false,
          "properties": {
            "nameFilter": {
              "description": "Optional data filename filter pattern that input files need to adhere to when type is 'data'. Eg parts of the expected filename",
              "type": "string"
            },
            "dataFormat": {
              "description": "Optional dataformat name array that input files need to adhere to when type is 'data'",
              "type": "array",
              "contains": {
                "type": "string"
              }
            },
            "dataType": {
              "description": "Optional data type (file or directory) that input files need to adhere to when type is 'data'",
              "type": "string",
              "enum": [
                "file",
                "directory"
              ]
            }
          }
        },
        "regex": {
          "type": "string"
        },
        "regexErrorMessage": {
          "type": "string"
        },
        "hidden": {
          "type": "boolean"
        },
        "disabled": {
          "type": "boolean"
        },
        "emptyValuesAllowed": {
          "type": "boolean",
          "description": "When maxValues is greater than 1 and emptyValuesAllowed is true, the values may contain null entries. Default is false."
        },
        "updateRenderOnChange": {
          "type": "boolean",
          "description": "When true, the onRender javascript function is triggered ech time the user changes the value of this field. Default is false."
        },
        "streamable": {
          "type": "boolean",
          "description": "EXPERIMENTAL PARAMETER! Only possible for fields of type 'data'. When true, the data input files will be offered in streaming mode to the pipeline instead of downloading them."
        },
      "required": [
        "id",
        "type"
      ],
      "allOf": [
        {
          "if": {
            "description": "When type is 'textbox' then 'dataFilter', 'fields', 'choices', 'max' and 'min' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "textbox"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "dataFilter",
                  "fields",
                  "choices",
                  "max",
                  "min"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'checkbox' then 'dataFilter', 'fields', 'choices', 'placeHolderText', 'regex', 'regexErrorMessage', 'maxLength', 'minLength', 'max' and 'min' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "checkbox"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "dataFilter",
                  "fields",
                  "choices",
                  "placeHolderText",
                  "regex",
                  "regexErrorMessage",
                  "maxLength",
                  "minLength",
                  "max",
                  "min"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'radio' then 'dataFilter', 'fields', 'placeHolderText', 'regex', 'regexErrorMessage', 'maxLength', 'minLength', 'max' and 'min' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "radio"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "dataFilter",
                  "fields",
                  "placeHolderText",
                  "regex",
                  "regexErrorMessage",
                  "maxLength",
                  "minLength",
                  "max",
                  "min"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'select' then 'dataFilter', 'fields', 'regex', 'regexErrorMessage', 'maxLength', 'minLength', 'max' and 'min' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "select"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "dataFilter",
                  "fields",
                  "regex",
                  "regexErrorMessage",
                  "maxLength",
                  "minLength",
                  "max",
                  "min"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'number' or 'integer' then 'dataFilter', 'fields', 'choices', 'regex', 'regexErrorMessage', 'maxLength' and 'minLength' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "number",
                  "integer"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "dataFilter",
                  "fields",
                  "choices",
                  "regex",
                  "regexErrorMessage",
                  "maxLength",
                  "minLength"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'data' then 'dataFilter' is required and 'fields', 'choices', 'placeHolderText', 'regex', 'regexErrorMessage', 'maxLength', 'minLength', 'max' and 'min' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "data"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "required": [
              "dataFilter"
            ],
            "propertyNames": {
              "not": {
                "enum": [
                  "fields",
                  "choices",
                  "placeHolderText",
                  "regex",
                  "regexErrorMessage",
                  "max",
                  "min",
                  "maxLength",
                  "minLength"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'section' or 'text' then 'disabled', 'fields', 'updateRenderOnChange', 'classification', 'value', 'minValues', 'maxValues', 'minMaxValuesMessage', 'dataFilter', 'choices', 'placeHolderText', 'regex', 'regexErrorMessage', 'maxLength', 'minLength', 'max' and 'min' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "section",
                  "text"
                ]
              }
            },
            "required": [
              "type"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "disabled",
                  "fields",
                  "updateRenderOnChange",
                  "classification",
                  "value",
                  "minValues",
                  "maxValues",
                  "minMaxValuesMessage",
                  "dataFilter",
                  "choices",
                  "regex",
                  "placeHolderText",
                  "regexErrorMessage",
                  "maxLength",
                  "minLength",
                  "max",
                  "min"
                ]
              }
            }
          }
        },
        {
          "if": {
            "description": "When type is 'fieldgroup' then 'fields' is required and then 'dataFilter', 'choices', 'placeHolderText', 'regex', 'regexErrorMessage', 'maxLength', 'minLength', 'max' and 'min' and 'emptyValuesAllowed' are not allowed",
            "properties": {
              "type": {
                "enum": [
                  "fieldgroup"
                ]
              }
            },
            "required": [
              "type",
              "fields"
            ]
          },
          "then": {
            "propertyNames": {
              "not": {
                "enum": [
                  "dataFilter",
                  "choices",
                  "placeHolderText",
                  "regex",
                  "regexErrorMessage",
                  "maxLength",
                  "minLength",
                  "max",
                  "min",
                  "emptyValuesAllowed"
                ]
              }
            }
          }
        }
      ]
    },
    "ica_pipeline_input_form_field_choice": {
      "$id": "#ica_pipeline_input_form_field_choice",
      "type": "object",
      "additionalProperties": false,
      "properties": {
        "value": {
        "description": "The value which will be set when selecting this choice. Must be unique over the choices within a field"
        },
        "text": {
          "description": "The display text for this choice, similar as the label of a field. ",
          "type": "string"
        },
        "selected": {
          "description": "Optional. When true, this choice value is picked as default selected value.  As in selected=true has precedence over an eventual set field 'value'. For clarity it's better however not to use 'selected' but use field 'value' as is used to set default values for the other field types.  Only maximum 1 choice may have selected true.",
          "type": "boolean"
        },
        "disabled": {
          "type": "boolean"
        },
        "parent": {
          "description": "Value of the parent choice item. Can be used to build hierarchical choice trees."
        }
      },
      "required": [
        "value",
        "text"
      ]
    }
  }
}
}
```

{% endcode %}


# onRender.js

Receives an input object which contains information about the current state of the input form, the chosen values and the field value change that triggered the onrender call. It also contains pipeline information. Changed objects are present in the onRender return value object. Any object not present is considered to be unmodified. Changing the storage size in the start analysis screen triggers an onRender execution with storageSize as changed field.

#### Input Parameters

<table><thead><tr><th width="235"></th><th></th></tr></thead><tbody><tr><td>context</td><td><p>"Initial"/"FieldChanged"/"Edited".</p><ul><li><strong>Initial</strong> is the value when first displaying the form when a user opens the <em>start run</em> screen.</li><li>The value is <strong>FieldChanged</strong> when a field with <code>'updateRenderOnChange'=true</code> is changed by the user.</li><li><strong>Edited</strong> (Not yet supported in Platform Core) is used when a form is displayed later again, this is intended for draft runs or when editing the form during reruns.</li></ul></td></tr><tr><td>changedFieldId</td><td>The id of the field that changed and which triggered this onRender call. context will be <code>FieldChanged</code>. When the storage size is changed, the fieldId will be <code>storageSize</code>.</td></tr><tr><td>analysisSettings</td><td>The input form json as saved in the pipeline. This is the original json, without changes.</td></tr><tr><td>currentAnalysisSettings</td><td>The current input form json as rendered to the user. This can contain already applied changes form earlier onRender passes. Null in the first call, when context is <code>Initial</code>.</td></tr><tr><td>settingValues</td><td>The current value of all settings fields. This is a map with field id as key and an array of field values as value for multivalue fields. For convenience, values of single-value fields are present as the individual value and not as an array of length 1. In case of fieldGroups, the value can be multiple levels of arrays. For fields of type <em>data</em> the values in the json are data ids (fil.xxxx). To help with validation, these are expanded and made available as an object here containing the <em>id</em>, <em>name</em>, <em>path</em>, <em>format</em>, <em>size</em> and a boolean indicating whether the data is <em>external</em>. This info can be used to validate or pick the chosen storageSize.</td></tr><tr><td>pipeline</td><td>Information about the pipeline: code, tenant and description are all available in the pipeline object as string.</td></tr><tr><td>analysis</td><td>Information about this run: userReference, userName, and userTenant are all available in the analysis object as string.</td></tr><tr><td>storageSize</td><td>The storage size as chosen by the user. This will initially be null. StorageSize is an object containing an 'id' and 'name' property.</td></tr><tr><td>storageSizeOptions</td><td>The list of storage sizes available to the user when creating an analysis. Is a list of StorageSize objects containing an 'id' and 'name' property.</td></tr></tbody></table>

#### Return values (taken from the response object)

<table><thead><tr><th width="191">Value</th><th>Meaning</th></tr></thead><tbody><tr><td>analysisSettings</td><td>The input form json with potential applied changes. The discovered changes will be applied in the UI.</td></tr><tr><td>settingValues</td><td>The current, potentially altered map of all setting values. These will be updated in the UI.</td></tr><tr><td>validationErrors</td><td>A list of RenderMessages representing validation errors. <strong>Submitting a pipeline execution request is not possible while there are still validation errors.</strong></td></tr><tr><td>validationWarnings</td><td>A list of RenderMessages representing validation warnings. <strong>A user may choose to ignore these validation warnings and start the pipeline execution request.</strong></td></tr><tr><td>storageSize</td><td><p>The suitable value for storageSize. Must be one of the options of input.storageSizeOptions. When absent or null, it is ignored.</p><p>validation errors and validation warnings can use 'storageSize' as fieldId to let an error appear on the storage size field. 'storageSize' is the value of the changedFieldId when the user alters the chosen storage size.</p></td></tr></tbody></table>

#### RenderMessage

This is the object used for **representing validation errors and warnings**. The attributes can be used with first letter lowercase (consistent with the input object attributes) or uppercase.

<table><thead><tr><th width="204">Value</th><th>Meaning</th></tr></thead><tbody><tr><td>fieldId / FieldId</td><td>The field which has an erroneous value. When not present, a general error/warning is displayed. To display an error on the storage size, use the <code>storageSize</code>Fieldid.</td></tr><tr><td>index / Index</td><td>The 0-starting index of the value which is incorrect. Use this when a particular value of a multivalue field is not correct. When not present, the entire field is marked as erroneous. The value can also be an array of indexes for use with fieldgroups. For instance, when the 3rd field of the 2nd instance of a fieldgroup is erroneous, a value of [ 1 , 2 ] is used.</td></tr><tr><td>message / Message</td><td>The error/warning message to display.</td></tr></tbody></table>


# onSubmit.js

The onSubmit.js javascript function receives an input object which holds information about the chosen values of the input form and the pipeline and pipeline execution request parameters. This javascript function is not only triggered when submitting a new pipeline execution request in the user interface, but also when submitting one through the rest API..

#### Input parameters

<table><thead><tr><th width="210">Value</th><th>Meaning</th></tr></thead><tbody><tr><td>settings</td><td>The value of the setting fields. Corresponds to <code>settingValues</code> in the onRender.js. This is a map with field id as key and an array of field values as value. For convenience, values of single-value fields are present as the individual value and not as an array of length 1. In case of fieldGroups, the value can be multiple levels of arrays. For fields of type <em>data</em> the values in the json are data ids (fil.xxxx). To help with validation, these are expanded and made available as an object here containing the <em>id</em>, <em>name</em>, <em>path</em>, <em>format</em>, <em>size</em> and a boolean indicating whether the data is <em>external</em>. This info can be used to validate or pick the chosen storageSize.</td></tr><tr><td>settingValues</td><td>To maximize the opportunity for reusing code between onRender and onSubmit, the 'settings' are also exposed as <code>settingValues</code> like in the onRender input.</td></tr><tr><td>pipeline</td><td>Info about the pipeline: code, tenant, and description are all available in the pipeline object as string.</td></tr><tr><td>analysis</td><td>Info about this run: userReference, userName, and userTenant are all available in the analysis object as string.</td></tr><tr><td>storageSize</td><td>The storage size as chosen by the user. This will initially be null. StorageSize is an object containing an 'id' and 'name' property.</td></tr><tr><td>storageSizeOptions</td><td>The list of storage sizes available to the user when creating an analysis. Is a list of StorageSize objects containing an 'id' and 'name' property.</td></tr><tr><td>analysisSettings</td><td>The input form json as saved in the pipeline. So the original json, without eventual changes.</td></tr><tr><td>currentAnalysisSettings</td><td>The current input form JSON as rendered to the user. This can contain already applied changes form earlier onRender passes. Null in the first call, when context is 'Initial' or when analysis is created through CLI/API.</td></tr></tbody></table>

#### Return values (taken from the response object)

<table><thead><tr><th width="191">Value</th><th>Meaning</th></tr></thead><tbody><tr><td>settings</td><td>The value of the setting fields. This allows modifying the values or applying defaults and such. Or taking info of the pipeline or analysis input object. When settings are not present in the onSubmit return value object, they are assumed to be not modified.</td></tr><tr><td>validationErrors</td><td>A list of AnalysisError essages representing validation errors. <strong>Submitting a pipeline execution request is not possible while there are still validation errors.</strong></td></tr><tr><td>analysisSettings</td><td>The input form json with potential applied changes. The discovered changes will be applied in the UI when viewing the analysis.</td></tr></tbody></table>

#### AnalysisError

This is the object used for representing validation errors.

<table><thead><tr><th width="204">Value</th><th>Meaning</th></tr></thead><tbody><tr><td>fieldId / FieldId</td><td>The field which has an erroneous value. When not present, a general error/warning is displayed. To display an error on the storage size, use the <code>storageSize</code>Fieldid.</td></tr><tr><td>index / Index</td><td>The 0-starting index of the value which is incorrect. Use this when a particular value of a multivalue field is not correct. When not present, the entire field is marked as erroneous. The value can also be an array of indexes for use with fieldgroups. For instance, when the 3rd field of the 2nd instance of a fieldgroup is erroneous, a value of [ 1 , 2 ] is used.</td></tr><tr><td>message / Message</td><td>The error/warning message to display.</td></tr></tbody></table>


# onFinished.js

The **onFinished** JavaScript functionality allows you to execute custom JavaScript code automatically after a pipeline analysis completes. It runs when the analysis completes successfully and provides access to the analysis inputs, outputs, and pipeline settings through a JavaScript execution environment. You can use this to perform post-processing tasks such as linking output files to samples.

## Usage <a href="#onfinishjavascript-gettingstarted" id="onfinishjavascript-gettingstarted"></a>

### Creating an onFinish Script <a href="#onfinishjavascript-creatinganonfinishscript" id="onfinishjavascript-creatinganonfinishscript"></a>

1. Create a JavaScript file named `onFinished.js` in your pipeline.
2. Define a function called `onFinished` that accepts an `input` parameter
3. The function will be called automatically when your analysis completes

## Basic Structure <a href="#onfinishjavascript-basicstructure" id="onfinishjavascript-basicstructure"></a>

```js
function onFinished(input) {
    // Your post-processing logic here
    print("Analysis completed!");
    
    // Access analysis information
    print("Pipeline: " + input.pipeline.name);
    print("Analysis status: " + input.analysis.status);
}
```

## Input Object <a href="#onfinishjavascript-inputobject" id="onfinishjavascript-inputobject"></a>

The `input` parameter provides access to analysis context and results:

```
input.pipeline
```

Information about the pipeline that was executed:

* **name** - Pipeline name
* **version** - Pipeline version string
* **tenant** - Tenant name
* **description** - Pipeline description

```
input.analysis
```

Information about the analysis execution:

* **userReference** - User-defined reference/name for the analysis
* **owner** - Username of the analysis owner
* **tenant** - Tenant name
* **status** - Current analysis status

```
input.currentAnalysisSettings
```

The pipeline input form configuration used for this analysis.

```
input.settingValues
```

A map of the actual input values provided when launching the analysis. Keys correspond to pipeline input parameter codes.

Example:

```js
// Access input parameter values
var sampleId = input.settingValues.sample_id;
var threshold = input.settingValues.quality_threshold;
```

```
input.outputs
```

A map of pipeline output parameters to their resulting data files. Keys are output parameter codes, values are arrays of data objects.

Each data object contains:

* **id** - Data file identifier
* **name** - Data file name
* **path** - Full path to the data file
* **format** - Data format code (e.g., "BAM", "FASTQ", "VCF")
* **size** - File size in bytes (-1 if unknown)
* **folder** - Boolean indicating if this is a folder
* **sampleIds** - Array of sample UUIDs associated with this data

```
input.samples
```

A map of sample objects referenced in the analysis inputs. Keys in the map are sample UUID, and every sample object contains:

* **id** - the UUID identifier
* **name** - Data file name
* **instrumentRunIds** - list of instrument run ids

## Injected Functions <a href="#onfinishjavascript-injectedfunctions" id="onfinishjavascript-injectedfunctions"></a>

The onFinish environment provides a number functions ([searchSamples](#onfinishjavascript-1.searchsamples), [searchData](#onfinishjavascript-2.searchdatas) and [linkDataToSamples](#onfinishjavascript-3.linkdatatosample)) for post-processing operations:

### searchSamples() <a href="#onfinishjavascript-1.searchsamples" id="onfinishjavascript-1.searchsamples"></a>

Search for samples by ID or name.

#### **Syntax:**

{% code overflow="wrap" expandable="true" %}

```js
var samples = searchSamples({
    id: "sample-uuid",        // Optional: Search by sample UUID
    name: "sample-name"       // Optional: Search by exact sample name
});
```

{% endcode %}

#### **Parameters:**

* Object with one or both (at least one parameter must be provided) of:
  * **id** - Sample UUID to search for
  * **name** - Exact sample name to search for

#### **Returns:**

Array of sample objects, where each sample contains:

* **id** - Sample UUID
* **name** - Sample name
* **instrumentRunIds** - Array of instrument run IDs associated with this sample

#### **Example:**

{% code overflow="wrap" expandable="true" %}

```js
// Search by sample ID
var samples = searchSamples({ id: "550e8400-e29b-41d4-a716-446655440000" });

// Search by sample name
var samples = searchSamples({ name: "SAMPLE_001" });

// Check results
if (samples && samples.length > 0) {
    samples.forEach(function(sample) {
        print("Found sample: " + sample.name + " (ID: " + sample.id + ")");
    });
} else {
    print("No samples found");
}
```

{% endcode %}

#### **Error Handling:**

Errors are automatically printed to the logs. The function returns null on error.

***

### searchDatas() <a href="#onfinishjavascript-2.searchdatas" id="onfinishjavascript-2.searchdatas"></a>

Search for data files in your analysis outputs with flexible filtering options.

#### **Syntax:**

{% code overflow="wrap" expandable="true" %}

```js
var datas = searchDatas({
    parent: dataObject,       // Optional: Parent folder/data to search within
    recursive: true,          // Optional: Search recursively (default: true)
    pathRegex: ".*\\.bam",  // Optional: Regex to match file paths
    nameRegex: ".*output.*"   // Optional: Regex to match file names
});
```

{% endcode %}

#### **Parameters:**

Object with optional filters:

* **parent** - Data object to use as parent (searches within this folder/path)
* **recursive** - Boolean, if **true** searches recursively from parent path (default: true)
* **pathRegex** - Regular expression to filter by file path (case-insensitive)
* **nameRegex** - Regular expression to filter by file name (case-insensitive)

#### **Returns:**

Array of data objects, where each data contains:

* **id** - Data file identifier
* **name** - File name
* **path** - Full file path
* **format** - Data format code
* **size** - File size in bytes
* **folder** - Boolean indicating if this is a folder
* **sampleIds** - Array of associated sample UUIDs

#### **Examples:**

{% code overflow="wrap" expandable="true" %}

```js
// Search for all BAM files in outputs
var bamFiles = searchDatas({
    nameRegex: ".*\\.bam$"
});

// Search within a specific output folder
var outputFolder = input.outputs.results[0];
var resultsFiles = searchDatas({
    parent: outputFolder,
    recursive: true
});

// Find files matching a pattern in a specific path
var reportFiles = searchDatas({
    pathRegex: ".*/reports/.*",
    nameRegex: ".*\\.html$"
});

// Process results
if (bamFiles && bamFiles.length > 0) {
    print("Found " + bamFiles.length + " BAM files:");
    bamFiles.forEach(function(file) {
        print("  - " + file.name + " at " + file.path);
    });
}
```

{% endcode %}

#### **Use Cases:**

* Find specific output files by pattern
* Navigate through output folder structures
* Filter files by type or naming convention
* Locate files for further processing

***

### linkDataToSample() <a href="#onfinishjavascript-3.linkdatatosample" id="onfinishjavascript-3.linkdatatosample"></a>

Link an output data file to a sample, establishing the relationship between analysis outputs and samples.

#### **Syntax:**

{% code overflow="wrap" expandable="true" %}

```js
linkDataToSample({
    data: dataObject,      // Required: Data object to link
    sample: sampleObject   // Required: Sample object to link to
});
```

{% endcode %}

#### **Parameters:**

Object with two required properties:

* **data** - Data object (from `input.outputs` or `searchDatas()`)
* **sample** - Sample object (from `input.samples` or `searchSamples()`)

#### **Returns:**

None. The function establishes the link in the system.

#### **Example:**

{% code overflow="wrap" expandable="true" %}

```js
// Link output BAM files to their corresponding samples
var outputBams = input.outputs.output_bam;

if (outputBams && outputBams.length > 0) {
    outputBams.forEach(function(bam) {
        // Extract sample name from BAM filename
        var sampleName = bam.name.replace(/\.bam$/, "");
        
        // Find the corresponding sample
        var samples = searchSamples({ name: sampleName });
        
        if (samples && samples.length > 0) {
            var sample = samples[0];
            
            // Link the BAM file to the sample
            linkDataToSample({
                data: bam,
                sample: sample
            });
            
            print("Linked " + bam.name + " to sample " + sample.name);
        } else {
            print("Warning: No sample found for " + bam.name);
        }
    });
}
```

{% endcode %}

#### **Use Cases:**

* Associate output files with input samples
* Track which files belong to which samples
* Enable sample-centric data organization
* Facilitate downstream sample-based queries

#### **Error Handling:**

If the data or sample cannot be found, or if the linking fails, an error message is automatically printed to the logs.

***

## Full Example <a href="#onfinishjavascript-completeexample" id="onfinishjavascript-completeexample"></a>

This example demonstrates a typical post-processing workflow:

{% code overflow="wrap" expandable="true" %}

```js
function onFinished(input) {
    print("=== Post-Analysis Processing Started ===");
    print("Pipeline: " + input.pipeline.name + " v" + input.pipeline.version);
    print("Analysis: " + input.analysis.userReference);
    
    // 1. Search for BAM files in the outputs
    print("\n--- Searching for BAM files ---");
    var bamFiles = searchDatas({
        nameRegex: ".*\\.bam$"
    });
    
    if (!bamFiles || bamFiles.length === 0) {
        print("No BAM files found in outputs");
        return;
    }
    
    print("Found " + bamFiles.length + " BAM file(s)");
    
    // 2. Link each BAM file to its corresponding sample
    print("\n--- Linking BAM files to samples ---");
    bamFiles.forEach(function(bam) {
        // Extract sample name from filename (assumes format: SAMPLENAME.bam)
        var fileBaseName = bam.name.replace(/\.bam$/, "");
        
        print("Processing: " + bam.name);
        
        // Search for the sample
        var samples = searchSamples({ name: fileBaseName });
        
        if (samples && samples.length > 0) {
            var sample = samples[0];
            
            // Create the link
            linkDataToSample({
                data: bam,
                sample: sample
            });
            
            print("  ✓ Successfully linked to sample: " + sample.name);
        } else {
            print("  ✗ Warning: No matching sample found for " + fileBaseName);
        }
    });
    
    // 3. Generate summary
    print("\n--- Summary ---");
    print("Total outputs processed: " + bamFiles.length);
   
    // List all output parameters
    for (var outputCode in input.outputs) {
        var outputData = input.outputs[outputCode];
        print("Output parameter '" + outputCode + "': " + outputData.length + " file(s)");
    }
    
    print("\n=== Post-Analysis Processing Completed ===");
}
```

{% endcode %}

#### Linking to Sample

link output files to the same sample to which the input file with the same name was linked.

{% code overflow="wrap" expandable="true" %}

```js
function onFinished(input) {
    print("hello, i'm inside the onFinished function");
    print("input = " + input);

    var inputFiles = input.settingValues["in"];

    var outputDatas = searchDatas({});
    print("outputDatas = " + outputDatas);

    for (const outputdata of outputDatas) {
        var inputFile = find(inputFiles, outputdata.name);
        if (inputFile && inputFile.sampleIds) {
            for (const sampleId of inputFile.sampleIds) {
                linkDataToSample({"data": outputdata, "sample": input.samples[sampleId]});
            }
        }
    }
}


function find(inputFiles, name) {
    for (const inputFile of inputFiles) {
        if (inputFile.name == name) {
            return inputFile;
        }
    }
    return null;
}
```

{% endcode %}

#### Linking with Textbox

Linking output files to a sample for which there is a textbox('targetSampleName') for the sample name in the inputform.

{% code expandable="true" %}

```js
function onFinished(input) {
    print("hello, i'm inside the onFinished function");
    print("input = " + input);
    var targetSampleName = input.settingValues["targetSampleName"];
    print("targetSampleName = " + targetSampleName);
    
    var validationErrors = [];

    var outputDatas = searchDatas({"nameRegex": ".*.txt"});
    print("outputDatas = " + outputDatas);
    
    var samples = searchSamples({"name": targetSampleName});
    print("samples = " + samples);
    
    for (const outputdata of outputDatas) {
        for (const sample of samples) {
            linkDataToSample({"data": outputdata, "sample": sample});
        }
    }
}
```

{% endcode %}

#### Linking Output Folders to Samples <a href="#onfinishjavascript-linkingoutputfolderstosamples" id="onfinishjavascript-linkingoutputfolderstosamples"></a>

{% code overflow="wrap" expandable="true" %}

```js
// When analysis outputs contain per-sample folders with multiple files
// Link each folder to its corresponding sample based on folder name

function onFinished(input) {
    print("=== Linking Sample Folders ===");
    
    var sampleFolders = getSampleOutputFolders(input);
    if (!sampleFolders || sampleFolders.length === 0) {
        print("No sample folders found in outputs");
        return;
    }
    
    print("Found " + sampleFolders.length + " output folder(s)");
    
    var successCount = 0;
    var failureCount = 0;
    
    sampleFolders.forEach(function(folder) {
        // Only process folders (not individual files)
        if (!folder.folder) {
            print("Skipping non-folder item: " + folder.name);
            return;
        }
        
        print("\nProcessing folder: " + folder.name);
        
        // Extract sample name from folder name
        // Assumes folder naming convention like: "SAMPLE_001_results" or "SAMPLE_001"
        var folderName = folder.name;
        
        // Method 1: Direct match (folder name is sample name)
        var sampleName = folderName;
        
        // Method 2: Extract from pattern (e.g., "SAMPLE_001_results" -> "SAMPLE_001")
        // Uncomment and adjust pattern as needed:
        // var match = folderName.match(/^(.+?)_results$/);
        // if (match) {
        //     sampleName = match[1];
        // }
        
        // Method 3: Remove suffix/prefix
        // sampleName = folderName.replace(/_results$/, "");
        
        print("  Looking for sample: " + sampleName);
        
        // Search for the sample
        var samples = searchSamples({ name: sampleName });
        
        if (samples && samples.length > 0) {
            var sample = samples[0];
            
            // Link the entire folder to the sample
            linkDataToSample({
                data: folder,
                sample: sample
            });
            
            print("  Successfully linked folder '" + folder.name + "' to sample '" + sample.name + "'");
            
            successCount++;
        } else {
            print("  Warning: No matching sample found for '" + sampleName + "'");
            failureCount++;
        }
    });
    
    // Summary
    print("\n=== Summary ===");
    print("Successfully linked: " + successCount + " folder(s)");
    print("Failed to link: " + failureCount + " folder(s)");
    print("Total processed: " + sampleFolders.length + " folder(s)");
}


function getSampleOutputFolders(input) {
var outputfolder = input.outputs["Output"][0];
print("outputfolder=" + outputfolder);
var firstLevelFilesAndFolders = searchDatas({
"parent": outputfolder,
"recursive": false
});

print("firstLevelFilesAndFolders=" + firstLevelFilesAndFolders);
var sampleOutputFolders = [];

firstLevelFilesAndFolders.forEach(function (firstLevelFileOrFolder) {
if (firstLevelFileOrFolder.folder == true && "reports" != firstLevelFileOrFolder.name) {
sampleOutputFolders.push(firstLevelFileOrFolder);
}
}
);

print("sampleOutputFolders=" + sampleOutputFolders);

return sampleOutputFolders;
}
```

{% endcode %}

## Best Practices <a href="#onfinishjavascript-bestpractices" id="onfinishjavascript-bestpractices"></a>

#### Error Handling <a href="#onfinishjavascript-1.errorhandling" id="onfinishjavascript-1.errorhandling"></a>

Check if results exist before processing:

{% code overflow="wrap" %}

```js
var samples = searchSamples({ name: sampleName });
if (samples && samples.length > 0) {
    // Process samples
} else {
    print("Warning: No samples found");
}
```

{% endcode %}

#### Logging <a href="#onfinishjavascript-2.logging" id="onfinishjavascript-2.logging"></a>

Use `print()` statements to provide visibility into your post-processing:

{% code overflow="wrap" %}

```js
print("Starting sample linking for " + bamFiles.length + " files");
```

{% endcode %}

#### Iterate Safely <a href="#onfinishjavascript-3.iteratesafely" id="onfinishjavascript-3.iteratesafely"></a>

Verify array contents before iteration:

{% code overflow="wrap" %}

```js
if (outputBams && outputBams.length > 0) {
    outputBams.forEach(function(bam) {
        // Process BAM file
    });
}
```

{% endcode %}

### Debugging <a href="#onfinishjavascript-debugging" id="onfinishjavascript-debugging"></a>

#### Viewing Logs <a href="#onfinishjavascript-viewinglogs" id="onfinishjavascript-viewinglogs"></a>

After the analysis completes, the onFinish execution logs are available in the analysis details screen to users with edit rights on the pipeline. All `print()` statements and error messages will appear in these logs.

### Advanced Use Cases <a href="#onfinishjavascript-advancedusecases" id="onfinishjavascript-advancedusecases"></a>

#### Conditional Linking Based on File Attributes <a href="#onfinishjavascript-conditionallinkingbasedonfileattributes" id="onfinishjavascript-conditionallinkingbasedonfileattributes"></a>

{% code overflow="wrap" %}

```js
// Only link BAM files larger than 1GB
var largeFiles = searchDatas({ nameRegex: ".*\\.bam$" })
    .filter(function(file) {
        return file.size > 1000000000; // 1GB in bytes
    });

print("Processing " + largeFiles.length + " large BAM files");
// ... continue with linking logic
```

{% endcode %}


# JSON Scatter Gather Pipeline

Let's create the [Nextflow Scatter Gather pipeline ](/tutorials/nextflow/scatter-gather-nextflow)with a JSON input form.

{% hint style="info" %}
Pay close attention to uppercase and lowercase characters when creating pipelines.
{% endhint %}

Select **Projects > your\_project > Flow > Pipelines**. From the **Pipelines** view, click the **+Create > Nextflow** **> JSON based** button to start creating a Nextflow pipeline.

In the **Details** tab, add values for the required *Code* (unique pipeline name) and *Description* fields. *Nextflow Version* and *Storage size* defaults to preassigned values.

### Nextflow files

#### split.nf

First, we present the individual processes. Select **Nextflow files > + Create** and label the file **split.nf**. Copy and paste the following definition.

```groovy
process split {
    cpus 1
    memory '512 MB'
    
    input:
    path x
    
    output:
    path("split.*.tsv")
    
    """
    split -a10 -d -l3 --numeric-suffixes=1 --additional-suffix .tsv ${x} split.
    """
    }
```

#### sort.nf

Next, select **+Create** and name the file **sort.nf**. Copy and paste the following definition.

```groovy
process sort {
    cpus 1
    memory '512 MB'
    
    input:
    path x
    
    output:
    path '*.sorted.tsv'
    
    """
    sort -gk1,1 $x > ${x.baseName}.sorted.tsv
    """
}
```

#### merge.nf

Select **+Create** again and label the file **merge.nf**. Copy and paste the following definition.

<pre class="language-groovy"><code class="lang-groovy">process merge {
  cpus 1
  memory '512 MB'
 
  publishDir 'out', mode: 'move'
 
  input:
  path x
 
  output:
<strong>  path 'merged.tsv'
</strong> 
  """
<strong>  cat $x > merged.tsv
</strong>  """
}
</code></pre>

#### main.nf

Edit the main.nf file by navigating to the **Nextflow files > main.nf** tab and copying and pasting the following definition.

```groovy
nextflow.enable.dsl=2
 
include { sort } from './sort.nf'
include { split } from './split.nf'
include { merge } from './merge.nf'
 
params.myinput = "test.test"
 
workflow {
    input_ch = Channel.fromPath(params.myinput)
    split(input_ch)
    sort(split.out.flatten())
    merge(sort.out.collect())
}
```

Here, the operators *flatten* and *collect* are used to transform the emitting channels. The *Flatten* operator transforms a channel in such a way that every item of type Collection or Array is flattened so that each single entry is emitted separately by the resulting channel. The collect operator collects all the items emitted by a channel to a List and return the resulting object as a sole emission.

### Inputform files

On the Inputform files tab, edit the inputForm.json to allow selection of a file.

#### inputForm.json

```json
{
  "fields": [
    {
      "id": "myinput",
      "label": "myinput",
      "type": "data",
      "dataFilter": {
        "dataType": "file",
        "dataFormat": ["TSV"]
      },
      "maxValues": 1,
      "minValues": 1
    }
  ]
}
```

Click the Simulate button (at the bottom of the text editor) to preview the launch form fields.

The onSubmit.js and onRender.js can remain with their default scripts and are just shown here for reference.

#### onSubmit.js

```json
function onSubmit(input) {
    var validationErrors = [];

    return {
        'settings': input.settings,
        'validationErrors': validationErrors
    };
}
```

#### onRender.js

```json
function onRender(input) {

    var validationErrors = [];
    var validationWarnings = [];

    if (input.currentAnalysisSettings === null) {
        //null first time, to use it in the remainder of he javascript
        input.currentAnalysisSettings = input.analysisSettings;
    }

    switch(input.context) {
        case 'Initial': {
            renderInitial(input, validationErrors, validationWarnings);
            break;
        }
        case 'FieldChanged': {
            renderFieldChanged(input, validationErrors, validationWarnings);
            break;
        }
        case 'Edited': {
            renderEdited(input, validationErrors, validationWarnings);
            break;
        }
        default:
            return {};
    }

    return {
        'analysisSettings': input.currentAnalysisSettings,
        'settingValues': input.settingValues,
        'validationErrors': validationErrors,
        'validationWarnings': validationWarnings
    };
}

function renderInitial(input, validationErrors, validationWarnings) {
}

function renderEdited(input, validationErrors, validationWarnings) {
}

function renderFieldChanged(input, validationErrors, validationWarnings) {
}

function findField(input, fieldId){
    var fields = input.currentAnalysisSettings['fields'];
    for (var i = 0; i < fields.length; i++){
        if (fields[i].id === fieldId) {
            return fields[i];
        }
    }
    return null;
}
```

Click the `Save` button to save the changes.


# Git-Sourced Pipelines

## Introduction

**Git** can be used as source-control system for your pipelines. This offers portability, versioning and easy integration of existing public pipelines.

To use Git-sourced pipelines:

1. If you want to use a private repository, Create [Git credentials ](#git-credentials)with your [personal access tokens](#personal-access-token).
2. [Configure in Platform Core](#creating-the-pipeline-in-platform-core) **where on GitHub the pipeline is located**, which **tag or commit-id** of that pipeline to use and select the **credentials** if it is a private repository or to prevent rate limiting.

## Using Private Repositories

### Git Credentials

To access private pipelines stored on Git, you need to enter the Git credentials in Platform Core. This is done at **System Settings > Credentials > Create > Git Credential**. If you want to use a public repository, you do not need credentials or an access token. However, there is a limit to how many anonymous calls can be made to a GitHub repository, so if that limit is exceeded, you will encounter *repo lookup rate limit reached* and will not be able to import the pipeline. For this reason, **it is best practice to use Git credentials even for public repositories.**

Fill out the following fields:

* **Name**—Provide a name to easily identify your Git credentials
* **Git url**—This is fixed to <https://github.com>
* **Personal Access Token** —the personal access token. Either an existing one, or you can [generate](#personal-access-token) one on Git. When you save and reopen your credentials, the personal access token is hidden. If you want to make changes to the credentials, you will need to re-enter the personal access token. This is to prevent the access token from being copied.
* **Git username** — This field **appears after entering the personal access token** and is extracted from your token.

<figure><img src="/files/CitGz9cWqN33IfsQZmZj" alt="" width="563"><figcaption></figcaption></figure>

A link is provided to create your [personal access token](#personal-access-token) on [github.com](https://github.com/settings/tokens/new?description=Illumina%20Connected%20Analytics\&scopes=repo) in case you do not have a token.

{% hint style="info" %}
The link to the personal access token on GitHub.com will open a pop-up window, so if it does not open, you might have a pop-up blocker active.
{% endhint %}

You can **share the credentials** you own with other users of your tenant. To do so, select your credentials at **System Settings > Credentials** and choose **Share**. Those users will be able to use your credentials, but they will not be able to see the actual access token.

### Personal Access Token

Personal access tokens work in the same way as OAuth access tokens. You can create a new personal access token at <https://github.com/settings/tokens/>.

* The personal access token must have **scope:repo** (meaning all elements in the repo category must be selected).
* If you set an expiration date, all pipelines which use this personal access token will stop working on that date. For this reason, it is advised **not to set an expiration date unless absolutely necessary**.

Select **Generate token** at the bottom of the screen to get your personal access token. **Copy the value over to a secure location** after generating it.

## Creating the pipeline in Platform Core

Regardless of using a public or a private repository, you need to create the connection in Platform Core to pick up pipeline the from Git.

1. Create a new pipeline at **Projects > your\_project > flow > Pipelines > Create > Nextflow > from Git**.
2. Fill out the fields to access the pipeline. During configuration, a test will be performed to see if the repository can be reached with the provided credentials. *Even if the tests fail, you can still save the entered values, this is to prevent you from being blocked from creating a configuration if the repository is unavailable.*\
   If you enter the URL for a private repository before you have entered the credentials, Platform Core will try to access it as if it were a public GitHub URL until you have entered the credentials.

<figure><img src="/files/m7RGdJ3YgiJxcU69Lz78" alt="" width="375"><figcaption></figcaption></figure>

<table data-header-hidden><thead><tr><th width="151"></th><th></th></tr></thead><tbody><tr><td>Repository url (max 255 chars)</td><td>The url where the pipeline is located.<br>For example https://github.com/nf-core/demo<br>Do not use a / as last character of your path as this will result in a failed pipeline with <em>Remote resource not found</em> as error.</td></tr><tr><td>Pipeline name<br>(max 255 chars)</td><td>Name of your pipeline. This is automatically filled out when the repository URL has been validated, but you can change the name.</td></tr><tr><td>Version number</td><td>The version of your pipeline. This is not the same as the tag or commit-id, but used to identify the version of your pipeline to people using it. This field can hold 25 characters.</td></tr><tr><td>Git credential</td><td><strong>Only needed for private repositories</strong>. This is your previously configured <a href="#git-credentials-private-repository">Git credential</a>. If you need Git credentials, but you have not configured them in the previous step, you can use the create button to configure them now.</td></tr><tr><td>Main file path<br>(max 255 chars)</td><td>For NextFlow, If the <strong>main.nf</strong> is not in the root folder, edit the main file path to the main.nf file, otherwise keep main.nf.</td></tr><tr><td>Config file path<br>(max 255 chars)</td><td>When configuring a nextflow pipeline, this is usually <em>nextflow.config</em> and located in the root folder. Edit this field to match your equivalent file and folder.</td></tr><tr><td>Schema file path<br>(max 255 chars)</td><td>When configuring a nextflow pipeline, this is the file which describes the different parameters used by the workflow. This is usually <strong>nextflow_schema.json</strong>.</td></tr><tr><td>Version</td><td>Use the long <strong>commit-id</strong> or the tag to identify the version. Platform Core will automatically convert the tag to fill out the commit-id.</td></tr></tbody></table>

3. Select **Create** to start importing the pipeline. During import, the **inputForm.json** file is generated from the **nextflow\_schema.json** file.
   * While the pipeline is being imported, you can still choose **Abort import** by opening the pipeline details if you decide not to import the pipeline.
   * If the **import fails**, you can investigate the reason by opening the pipeline and looking at the **import summary on the details tab**.
   * After import, the pipeline will be in **draft** [status](/project/p-flow/f-pipelines#pipeline-statuses).
4. The details of your pipeline are now visible when opening the pipeline details tab. Here you can change the **name** and **version number**, provide a **description**, update the **Nextflow version** or edit other details as needed.
5. Proceed to the **InputForm files** **tab** to **edit the generated inputForm.json** file to match the inputs as needed by your pipeline. While you are editing the input form, you can use the **simulate** button to preview the input form.
6. Once all details are updated, **save your pipeline**. [Analyses](/project/p-flow/f-analyses) with the pipeline can now be executed in the same way as with regular pipelines.

{% hint style="info" %}
To minimise the risk of triggering API-rate limiting on GitHub, the validations above are not performed when configuring pipelines via the API.
{% endhint %}

## Example

{% content-ref url="/pages/E9aascoyvWbjDrN3zyEf" %}
[Basic Git-Sourced Pipeline Example](/project/p-flow/f-pipelines/git-sourced-pipelines-experimental/basic-git-sourced-pipeline-example)
{% endcontent-ref %}

## Restrictions

The following restrictions apply to Git-sourced pipelines.

* Only **JSON-based NextFlow** are supported, no CWL pipelines or XML-based pipelines.
* Because translation from the generic nextflow pipeline to the ica-specific input forms takes place during import, **you cannot use input forms created in Platform Core in your GitHub repository** as the system will try to reconvert them.
* **Only** [**Github.com**](http://github.com) is supported as repository location.
* GitHub enterprise is not supported.
* **Git-branches are not supported**, only tag or commit-id for version control.
* On GitHub, repository names are limited to 100 characters


# Basic Git-Sourced Pipeline Example

## Demo Pipeline

This section describes how to configure the **nf-core demo** pipeline from <https://github.com/nf-core/demo/> to run in Platform Core.

{% hint style="info" %}
**The nf-core framework for community-curated bioinformatics pipelines.**\
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.

*Nat Biotechnol.* 2020 Feb 13. doi: [10.1038/s41587-020-0439-x](https://dx.doi.org/10.1038/s41587-020-0439-x).
{% endhint %}

### Configuring the Pipeline

1. Create a new pipeline at **Projects > your\_project > flow > Pipelines > Create > Nextflow > from Git**.
2. Fill out the following details:

<table><thead><tr><th width="203.04296875">Field</th><th>Value</th></tr></thead><tbody><tr><td><strong>Repository url</strong></td><td><strong>https://github.com/nf-core/demo</strong> (Tip: If you encounter <em>Invalid GitHub repository url</em>, you may have copied over a trailing space in the url)</td></tr><tr><td><strong>Pipeline name</strong></td><td>The pipeline name is automatically extracted from the repository. You can manually change this if needed, for example to prevent duplicate names.</td></tr><tr><td><strong>Version number</strong></td><td>Enter the version number. You can choose to use the version number from GitHub (1.1.0) or you can use your own version (for example 1 for being the first local version)</td></tr><tr><td><strong>Storage size</strong></td><td>As this demo pipeline uses very little resources, the smallest size we can select (<strong>3XSmall)</strong> will be sufficient.</td></tr><tr><td><strong>Git credential</strong></td><td><p>This is a public repository, so we don't actually need a credential and this can be left blank. However, there is a limit to how many anonymous calls can be made to a GitHub repository, so if that limit is exceeded, you will encounter <em>repo lookup rate limit reached</em> and will not be able to import the pipeline.</p><ul><li>To prevent running into this limit, select a Git credential.</li><li>If you don't have one, click the <strong>create</strong> button next to the Git credential. Enter the value of your personal access token and Platform Core will automatically obtain the Git username for it.</li><li>If you have no access token, you can create one with the <strong>Create personal access token on github.com</strong> button. The values there are already prefilled, but you can change the note at the top to easily identify for which purpose the token was generated.</li></ul></td></tr><tr><td><strong>Main file path</strong></td><td>The main.nf file containing the pipeline logic and data flow is in the root folder, so we keep this as <strong>main.nf</strong>.</td></tr><tr><td><strong>Config file path</strong></td><td>nextflow.config containing the execution parameters is in the root folder, so we need to enter <strong>nextflow.config</strong>. The inputForm file will be generated based on the parameters in this file.</td></tr><tr><td><strong>Schema file path</strong></td><td>The schema is <strong>nextflow_schema.json</strong>, so enter this value. The inputForm file will be generated based on the parameters in this file.</td></tr><tr><td><strong>Version</strong></td><td><p>You can enter the <strong>commit id</strong> of the version you want to use or use the tag to identify the version.</p><p>In this example, we use the <strong>tags</strong> to identify the version we want. From the screenshot below, you can see that there is a version 1.1.0 of the pipeline, so enter <strong>1.1.0</strong> in the tag field. When you enter this version in Platform Core, the commit-id for that tag will automatically be filled out. This commit-id is the long version, <strong>the short 7-character version is not supported.</strong></p></td></tr></tbody></table>

<figure><img src="/files/EB44Geas5Hf8MhiGZMHB" alt=""><figcaption><p>Method 1 : finding and copying the full commit id</p></figcaption></figure>

<figure><img src="/files/w6iqraVqpsnZN3Hd3CEj" alt="" width="563"><figcaption><p>Method 2 : finding and copying the tag</p></figcaption></figure>

### Creating the Input File

As per [instructions](https://github.com/nf-core/demo/blob/master/README.md) on the nf-core demo page, **create a samplesheet.csv file** on your local machine with as contents:

```
sample,fastq_1,fastq_2
SAMPLE1_PE,https://raw.githubusercontent.com/nf-core/test-datasets/viralrecon/illumina/amplicon/sample1_R1.fastq.gz,https://raw.githubusercontent.com/nf-core/test-datasets/viralrecon/illumina/amplicon/sample1_R2.fastq.gz
SAMPLE2_PE,https://raw.githubusercontent.com/nf-core/test-datasets/viralrecon/illumina/amplicon/sample2_R1.fastq.gz,https://raw.githubusercontent.com/nf-core/test-datasets/viralrecon/illumina/amplicon/sample2_R2.fastq.gz
SAMPLE3_SE,https://raw.githubusercontent.com/nf-core/test-datasets/viralrecon/illumina/amplicon/sample1_R1.fastq.gz,
SAMPLE3_SE,https://raw.githubusercontent.com/nf-core/test-datasets/viralrecon/illumina/amplicon/sample2_R1.fastq.gz,
```

#### Using the GUI

1. In Platform Core, navigate to **Projects > your\_project > Data** and upload the created **samplesheet.csv** file to your project.

#### Using the CLI

If you have the [CLI](/command-line-interface/cli-indexcommands) installed on your system, you can use the commands below to upload the samplesheet. If you do not have an active CLI, please follow [these instructions](/command-line-interface/cli-installation) first.

1. List your projects with the command `icav2 projects list`
2. From this list of projects, enter your demo project by using `icav2 projects enter <your_project_uuid>` with your\_project\_uuid replaced with the uuid of your project.
3. If you have created the samplesheet.csv file and put it into your CLI directory, you upload it to the root folder of your project with `icav2 projectdata upload samplesheet.csv` . If this name is already in use, you can rename the file during upload by using `icav2 projectdata upload samplesheet.csv /samplesheet2.csv`

### Running the Analysis

#### Using the GUI

With the pipeline configurad and the inputfile created, you are ready to run your analysis.

1. Go to **Projects > your\_project > flow > Pipelines.**
2. **Select the created pipeline.** (The default name will be nf-core/demo) and choose **Start analysis** at the top of the screen.
3. You will be presented with the form below.
   * Enter an identifier (**user reference**) for your pipeline
   * Select the **samplesheet.csv** file as **input** and start the analysis.

<figure><img src="/files/oo9H5W3iw0peH7PpU6uo" alt=""><figcaption></figcaption></figure>

You can follow the status of your analysis at **Projects > your\_project > Flow > Analyses**.

{% hint style="info" %}
If your analysis fails with *Module path must start with / or ./ prefix -- Offending module: plugin/nf-schema*, then the Nextlow version is not set to the latest one. Edit it on your **Projects > your\_project > flow > Pipelines > your\_pipeline > Details**
{% endhint %}

#### Using the CLI

If you have the [CLI](/command-line-interface/cli-indexcommands) installed on your system, you can also use the commands below to work with pipelines. If you do not have an active CLI, please follow [these instructions](/command-line-interface/cli-installation) first.

1. List your projects with the command `icav2 projects list` to retrieve their **project uuid**.
2. From this list of projects, enter your demo project by using `icav2 projects enter <your_project_uuid>` with \<your\_project\_uuid> replaced with the uuid of your project.
3. To list the pipelines in your project and see their **pipeline uuid**, use `icav2 projectpipelines list`
4. To retrieve the **file uuid**, use the command `icav2 projectdata list --file-name samplesheet.csv` . If you used a different name for your samplesheet file, use the name you gave it instead of samplesheet.csv.

{% hint style="info" %}
You can search for csv files in your project with the command `icav2 projectdata list --file-name csv --match-mode fuzzy`.
{% endhint %}

5. To run the pipeline with the input file,

* Select the **pipeline** (replace \<your\_pipeline\_uuid> with the uuid of your pipeline).
* Set the **storage size** (3XSmall). If this storage is not available in your subscription, you can use `icav2 analysisstorages list` to get a list of available storage sizes.
* Set the **user reference** to give your analysis a name. In this example we use MyDemoGitPipeline as name.
* Point to the **input file** (replace \<your\_file\_uuid> with the actual file uuid).

  If you want to see a list of input parameters of your pipeline, use `icav2 projectpipelines input <your_pipeline_uuid>` this will show you the code of the input parameters. In our example, this is "input"

<figure><img src="/files/2njKx8VynSKIiUoUzhvv" alt="" width="563"><figcaption></figcaption></figure>

The resulting command to run the pipeline is then:

{% code overflow="wrap" %}

```bash
icav2 projectpipelines start nextflowjson <your_pipeline_uuid> --storage-size 3XSmall --user-reference MyDemoGitPipeline --field-data "input":"<your_file_uuid>"
```

{% endcode %}

### Optional : Editing the Input Form

Even though the import function will have created an input form based on the configured files, you can customise this form to make it easier to use by removing or defaulting parameters.

1. Go to **Projects > your\_project > flow > Pipelines > your\_pipeline > Edit**
2. Navigate to the **Inputform files** tab. This will open the inputForm.json file.
3. Replace the contents with this minimalist input file

```json
{
  "fields": [
    {
      "id": "input",
      "type": "data",
      "label": "input",
      "helpText": "Select your input file",
      },
      "maxValues": 1,
      "minValues": 1
    },
    {
      "id": "outdir",
      "type": "textbox",
      "label": "outdir",
      "helpText": "The output directory where the results will be saved.",
      "hidden": true,
      "value": "out",
      "minValues": 1
    }
  ]
}
```

This will result in a minimal input form which only needs the input file selection.

<figure><img src="/files/gy3xbELxP3SAC9GP7Qro" alt=""><figcaption></figcaption></figure>


# Tips and Tricks

Developing on the cloud incurs inherent runtime costs due to compute and storage used to execute workflows. Here are a few tips that can facilitate development:

* Leverage the cross-platform nature of these workflow languages. Both CWL and Nextflow can be run locally in addition to on Platform Core. When possible, testing should be performed locally before attempting to run in the cloud. For Nextflow, [configuration files](https://www.nextflow.io/docs/latest/config.html) can be utilized to specify settings to be used either locally or on Platform Core. An example of advanced usage of a config would be applying the [scratch directive](https://www.nextflow.io/docs/latest/process.html#scratch) to a set of process names (or labels) so that they use the higher performance local scratch storage attached to an instance instead of the shared network disk,

  ```groovy
  withName: 'process1|process2|process3' { scratch = '/scratch/' }
  withName: 'process3' { stageInMode = 'copy' } // Copy the input files to scratch instead of symlinking to shared network disk
  ```
* When trying to test on the cloud, it's oftentimes beneficial to create scripts to automate the deployment and launching / monitoring process. This can be performed either using the [CLI](broken://spaces/7GiJwg33pKa8eXmwcle7/pages/E6sDxxcHWmVagLfpsfEz) or by creating your own scripts integrating with the REST API.
* For scenarios in which instances are terminated prematurely (for example, while using spot instances) without warning, you can implement scripts like the following to retry the job a certain number of times. Adding the following script to 'nextflow\.config' enables five retries for each job, with increasing delays between each try.

  ```groovy
  process {
      maxRetries = 4
      errorStrategy = { sleep(task.attempt * 60000 as long); return'retry'} // Retry with increasing delay
  }
  ```

  Note: Adding the retry script where it is not needed might introduce additional delays.
* When hardening a **Nextflow** to handle resource shortages (for example exit code 2147483647), an immediate retry will in most circumstances fail because the resources have not yet been made available. It is best practice to use [Dynamic retry with backoff](https://www.nextflow.io/docs/latest/process.html#dynamic-retry-with-backoff) which has an increasing backoff delay, allowing the system time to provide the necessary resources.
* When publishing your **Nextflow** pipeline, make sure your have defined a container such as 'public.ecr.aws/lts/ubuntu:22.04' and are not using the default container 'ubuntu:latest'.
* To limit potential costs, there is a timeout of 96 hours: if the analysis does not complete within four days, it will go to a 'Failed' state. This time begins to count as soon as the input data is being downloaded. This takes place during the Platform Core 'Requested' step of the analysis, before going to 'In Progress'. In case parallel tasks are executed, running time is counted once. As an example, let's assume the initial period before being picked up for execution is 10 minutes and consists of the request, queueing and initializing. Then, the data download takes 20 minutes. Next, a task runs on a single node for 25 minutes, followed by 10 minutes of queue time. Finally, three tasks execute simultaneously, each of them taking 25, 28, and 30 minutes, respectively. Upon completion, this is followed by uploading the outputs for one minute. The overall analysis time is then 20 + 25 + 10 + 30 (as the longest task out of three) + 1 = 86 minutes:

|      **Analysis task**      |      request     |      queued      |     initializing    |      input download     |     single task    |        queue       |   parallel tasks   |     generating outputs    |     completed    |
| :-------------------------: | :--------------: | :--------------: | :-----------------: | :---------------------: | :----------------: | :----------------: | :----------------: | :-----------------------: | :--------------: |
|      **96 hour limit**      | 1m (not counted) | 7m (not counted) |   2m (not counted)  |           20m           |         25m        |         10m        |         30m        |             1m            |         -        |
| **Status in Platform Core** | status requested |   status queued  | status initializing | status preparing inputs | status in progress | status in progress | status in progress | status generating outputs | status succeeded |

If there are no available resources or your project priority is low, the time before download commences will be substantially longer.

* By default, Nextflow will not generate the trace report. If you want to enable generating the report, add the section below to your userNextflow\.config file.

{% code overflow="wrap" %}

```
trace.enabled = true
trace.file = '.ica/user/trace-report.txt'
trace.fields = 'task_id,hash,native_id,process,tag,name,status,exit,module,container,cpus,time,disk,memory,attempt,submit,start,complete,duration,realtime,queue,%cpu,%mem,rss,vmem,peak_rss,peak_vmem,rchar,wchar,syscr,syscw,read_bytes,write_bytes,vol_ctxt,inv_ctxt,env,workdir,script,scratch,error_action'
```

{% endcode %}

7. Useful Links
   * [Nextflow on Kubernetes: Best Practices](https://github.com/seqeralabs/nf-k8s-best-practices)
   * [The State of Kubernetes in Nextflow](https://www.nextflow.io/blog/2023/the-state-of-kubernetes-in-nextflow.html)


# Analyses

An Analysis is the execution of a pipeline.

## Starting Analyses

You can start an analysis from both the dedicated analysis screen or from the actual pipeline.

#### From Analyses

1. Navigate to **Projects > Your\_Project > Flow > Analyses**.
2. Select **Start**.
3. Select a single Pipeline.
4. Configure the[ analysis settings](/project/p-flow/f-pipelines#analysis-tab).
5. Select **Start Analysis**.
6. Refresh to see the analysis status. See [lifecycle](#Lifecycle) for more information on statuses.
7. If for some reason, you want to end the analysis before it can complete, select **Projects > Your\_Project > Flow > Analyses > Manage > Abort**. Refresh to see the status update.

#### From Pipelines or Pipeline details

1. Navigate to **Projects > \<Your\_Project> > Flow > Pipelines**
2. Select the pipeline you want to run or open the pipeline details of the pipeline which you want to run.
3. Select **Start Analysis**.
4. Configure [analysis settings](/project/p-flow/f-pipelines#analysis-tab).
5. Select **Start Analysis**.
6. View the analysis status on the Analyses page. See [lifecycle](#Lifecycle) for more information on statuses.
7. If for some reason, you want to end the analysis before it can complete, select **Manage > Abort** on the Analyses page.

#### Aborting Analyses

You can abort a running analysis from either the analysis overview (**Projects > your\_project > Flow > Analyses > your\_analysis > Manage > Abort**) or from the analysis details (**Projects > your\_project > Flow > Analyses > your\_analysis > Details tab > Abort**).

#### Rerunning Analyses

Once an analysis has been executed, you can rerun it with the same settings or choose to modify the parameters when rerunning. Modifying the parameters is possible on a per-analysis basis. When selecting multiple analyses at once, they will be executed with the original parameters. **Draft pipelines** are subject to updates and thus can result in a different outcome when rerunning. Platform Core will display a warning message to inform you of this when you try to rerun an analysis based on a draft pipeline.

{% hint style="info" %}
When rerunning multiple analysis at the same time (multiselect), the user reference can not be chosen and will be the original user reference (up to 231 characters), followed by \_rerun\_yyyy-MM-dd\_HHmmss.
{% endhint %}

When there is an **XML configuration change** on a a pipeline for which you want to rerun an analysis, Platform Core will display a warning and not fill out the parameters as it cannot guarantee their validity for the new XML. You will need to provide the input data and settings again to rerun the analysis.

Some restrictions apply when trying to rerun an analysis.

<table><thead><tr><th width="374">Analyses</th><th width="138">Rerun</th><th>Rerun with modified parameters</th></tr></thead><tbody><tr><td>Analyses using external data</td><td>Allowed</td><td>-</td></tr><tr><td>Analyses using mount paths on input data</td><td>Allowed</td><td>-</td></tr><tr><td>Analyses using user-provided input json</td><td>Allowed</td><td>-</td></tr><tr><td>Analyses using advanced output mappings</td><td>-</td><td>-</td></tr><tr><td>Analyses with draft pipeline</td><td>Warn</td><td>Warn</td></tr><tr><td>Analyses with XML configuration change</td><td>Warn</td><td>Warn</td></tr></tbody></table>

To rerun one or more analyses with te same settings:

1. Navigate to **Projects > Your\_Project > Flow > Analyses**.
2. In the overview screen, **select one or more** analyses.
3. Select **Manage > Rerun**. The analyses will now be executed with the same parameters as their original run.

To rerun a single analysis with modified parameters:

1. Navigate to **Projects > Your\_Project > Flow > Analyses**.
2. In the overview screen, **open the details** of the analysis you want to rerun by clicking on the analysis user reference.
3. Select **Rerun**. (at the top right)
4. Update the parameters you want to change.
5. Select **Start Analysis** The analysis will now be executed with the updated parameters.

{% hint style="info" %}
You might see parameters (files and folders) on the analysis details tab which are not defined on the XML. This indicates they are added in the xml-pipeline analysis creation endpoint.
{% endhint %}

## Lifecycle <a href="#lifecycle" id="lifecycle"></a>

<table><thead><tr><th>Status</th><th width="393">Description</th><th>Final State</th></tr></thead><tbody><tr><td>Requested</td><td>The request to start the Analysis is being processed</td><td>No</td></tr><tr><td>Queued</td><td>Analysis has been queued</td><td>No</td></tr><tr><td>Initializing</td><td>Initializing environment and performing validations for Analysis</td><td>No</td></tr><tr><td>Preparing Inputs</td><td>Downloading inputs for Analysis</td><td>No</td></tr><tr><td>In Progress</td><td>Analysis execution is in progress</td><td>No</td></tr><tr><td>Generating outputs</td><td>Transferring the Analysis results</td><td>No</td></tr><tr><td>Aborting</td><td>Analysis has been requested to be aborted</td><td>No</td></tr><tr><td>Aborted</td><td>Analysis has been aborted</td><td>Yes</td></tr><tr><td>Failed</td><td>Analysis has finished with error</td><td>Yes</td></tr><tr><td>Succeeded</td><td>Analysis has finished with success</td><td>Yes</td></tr></tbody></table>

{% hint style="info" %}
During analysis start, Platform Core runs a verification on the input files to see if they are available. When it encounters files that have not completed their upload or transfer, it will report "*Data found for parameter \[parameter\_name], but status is Partial instead of Available*". Wait for the file to be available and restart the analysis.
{% endhint %}

{% hint style="info" %}
When an analysis is started, the availability of resources may impact the start time of the pipeline or specific steps after execution has started. Analyses are subject to delay when cloud resources are under high load.\
\
When the underlying storage provider runs out of storage resources, the Status field of the Analysis details will indicate this. There is no need to abort or rerun the analysis.
{% endhint %}

## Analysis steps logs

During the execution of an analysis, logs are produced for each process involved in the analysis lifecyle. In the analysis details view, the **Steps tab** is used to view the steps in near real time as they're produced in the running processes. A grid layout is used for analyses with more than 50 steps, a tiled view for analyses with 50 steps or less, though you can choose to also use the grid layout for those by means of the *tile/grid button* on the top right of the analysis log tab. The steps tab also shows **which resources were used** as compute type in the different main analysis steps. (For child steps, these are displayed on the parent step)

<figure><img src="/files/l71oKXMQqA5CoE65Ixoj" alt=""><figcaption></figcaption></figure>

There are system processes involved in the lifeycle for all analyses (ie. downloading inputs, uploading outputs, etc.) and there are processes which are pipeline-specific, such as processes which execute the pipeline steps. The below table describes the system processes. You can choose to display or hide these system processes with the Show technical steps

<table><thead><tr><th width="241">Process</th><th>Description</th></tr></thead><tbody><tr><td>Setup Environment</td><td>Validate analysis execution environment is prepared</td></tr><tr><td>Run Monitor</td><td>Monitor resource usage for billing and reporting</td></tr><tr><td>Prepare Input Data</td><td>Download and mount input data to the shared file system</td></tr><tr><td>Pipeline Runner</td><td>Parent process to execute the pipeline definition</td></tr><tr><td>Finalize Output Data</td><td>Upload Output Data</td></tr></tbody></table>

Additional log entries will show for the processes which execute the steps defined in the pipeline.

Each process shows as a distinct entry in the steps view with a Queue Date, Start Date, and End Date.

<table><thead><tr><th width="175">Timestamp</th><th>Description</th></tr></thead><tbody><tr><td>Queue Date</td><td>The time when the process is submitted to the processes scheduler for execution</td></tr><tr><td>Start Date</td><td>The time when the process has started exection</td></tr><tr><td>End Date</td><td>The time when the process has stopped execution</td></tr></tbody></table>

The time between the Start Date and the End Date is used to calculate the duration. The time of the duration is used to calculate the usage-based cost for the analysis. Because this is an active calculation, sorting on this field is not supported.

Each log entry in the Steps view contains a checkbox to view the stdout and stderr log files for the process. Clicking a checkbox adds the log as a tab to the log viewer where the log text is displayed and made available for download.

### Analysis Cost

To see the price of an analysis in BioInsight Credits (BIC), look at **Projects > your\_project > Flow > Analyses > your\_analysis > Details tab**. The pricing section will show you the entitlement bundle, storage detail and price in BIC once the analysis has succeeded, failed or been aborted.

### Log Files

By default, the **stdout** and **stderr** files are located in the ***ica\_logs*** subfolder within the analysis. This **location can be changed** by selecting a different [logs folder ](/project/p-flow/f-pipelines#analysis-settings)in the current project at the start of the analysis. **Do not use a folder which already contains log files** as these will be overwritten.\
To set the log file location, you can also use the CreateAnalysisLogs section of the Create Analysis [endpoints](https://ica.illumina.com/ica/api/swagger/index.html).

{% hint style="warning" %}
If you delete these files, no log information will be available on the **analysis details > Steps tab**.
{% endhint %}

You can access the log files from the analysis details (**projects > your\_project > flow > analysis > your\_analysis > details tab**)

### Log Streaming

Logs can also be streamed using websocket client tooling. The API to retrieve analysis step details returns websocket URLs for each step to stream the logs from stdout/stderr during the step's execution. Upon completion, the websocket URL is no longer available.

## Analysis Output Mappings

{% hint style="warning" %}
Currently, only FOLDER type output mappings are supported
{% endhint %}

By default, analysis outputs are directed to a new folder within the project where the analysis is launched. Analysis output mappings may be specified to redirect outputs to user-specified locations consisting of project and path. An output mapping consists of:

* the source path on the local disk of the analysis execution environment, relative to the working folder.
* the data type, either FILE or FOLDER
* the target project ID to direct outputs to; analysis launcher must have contributor access to the project.
* the target path relative to the root of the project data to write the outputs.

{% hint style="warning" %}
If the output folder already exists, any existing contents with the same filenames as those output from the pipeline will be overwritten by the new analysis
{% endhint %}

<details>

<summary>Example</summary>

In this example, 2 analysis output mappings are specified. The analysis writes data during execution in the working directory at paths `out/test` and `out/test2`. The data contained in these folders are directed to project with ID `4d350d0f-88d8-4640-886d-5b8a23de7d81` and at paths `/output-testing-01/` and `/output-testing-02/` respectively, relative to the root of the project data.

The following demonstrates the construction of the request body to start an analysis with the output mappings described above:

````
```json
{
...
    "analysisOutput":
    [
        {
            "sourcePath": "out/test1",
            "type": "FOLDER",
            "targetProjectId": "4d350d0f-88d8-4640-886d-5b8a23de7d81",
            "targetPath": "/output-testing-01/"
        },
        {
            "sourcePath": "out/test2",
            "type": "FOLDER",
            "targetProjectId": "4d350d0f-88d8-4640-886d-5b8a23de7d81",
            "targetPath": "/output-testing-02/"
        }
    ]
}
```
````

When the analysis completes the outputs can be seen in the Platform Core UI, within the folders designated in the payload JSON during pipeline launch (`output-testing-01` and `output-testing-02`).

</details>

You can jump from the Analysis Details to the individual files and folders by opening the output files tab on the detail view (P**rojects > your\_project > Flow > Analyses > your\_analysis > Output files tab > your\_output\_file**) and selecting O**pen in data**.

{% hint style="info" %}
The **Output files** section of the analyses will always show the generated outputs, even when they have since been deleted from storage. This is done so you can always see which files were generated during the analysis.\
In this case it will no longer be possible to navigate to the actual output files.
{% endhint %}

<table><thead><tr><th width="150.98046875">analysis output</th><th width="120.3671875">logs output</th><th>Notes</th></tr></thead><tbody><tr><td>Default</td><td>Default</td><td>Logs are a subfolder of the analysis output.</td></tr><tr><td>Mapped</td><td>Default</td><td>Logs are a subfolder of the analysis output.</td></tr><tr><td>Default</td><td>Mapped</td><td>Outputs and logs may be separated.</td></tr><tr><td>Mapped</td><td>Mapped</td><td>Outputs and logs may be separated.</td></tr></tbody></table>

## Tags

You can add and remove tags from your analyses.

1. Navigate to **Projects > Your\_Project > Flow > Analyses**.
2. Select the analyses whose tags you want to change.
3. Select **Manage > Manage tags**.
4. Edit the user tags, reference data tags (if applicable) and technical tags.
5. Select **Save** to confirm the changes.

Both system tags and customs tags exist. **User** tags are custom tags which you set to help identify and process information while **technical** tags are set by the system for processing. Both **run-in** and **run-out** tags are set on data to identify which analyses use the data. **Connector** tags determine data entry methods and **reference data** tags identify where data is used as reference data.

## Hyperlinking

If you want to share a link to an analysis, you can copy and paste the URL from your browser when you have the analysis open. The syntax of the analysis link will be `<hostURL>/ica/link/project/<projectUUID>/analysis/<analysisUUID>`. Likewise, workflow sessions will use the syntax `<hostURL>/ica/link/project/<projectUUID>/workflowSession/<workflowsessionUUID>`. To prevent third parties from accessing data via the link when it is shared or forwarded, Platform Core will verify the access rights of every user when they open the link.

## Restrictions

Input for analysis is limited to a total of 50,000 files (including multiple copies of the same file). Concurrency limits on analyses prevent resource hogging which could result in resource starvation for other tenants. Additional analyses will be queued and scheduled when currently running analyses complete and free up positions. The theoretical limit is 20, but this can be less in practice, depending on a number of external factors.

## Troubleshooting

When your analysis fails, open the analysis details view (**Projects > your\_project> Flow > Analyses > your\_analysis**) and select **display failed steps**. This will give you the steps view filtered on those steps that had non-0 exit codes. If there is only one failed step which has logfiles, the stderr of that step will be displayed.

{% hint style="info" %}
For pipeline developers: **add automatic retrying to the individual steps** that fail with error 55 / 56 (provided these steps are idempotent) See [tips and tricks](/project/p-flow/f-pipelines/pi-tips) for retries.
{% endhint %}

* Exit **code 55** indicates analysis failure on economy instances due to an external event such as spot termination. **You can retry the analysis.**
* Exit **code 56** indicates analysis failure due to pod disruption and deletion by Kubernetes' Pod Garbage Collector (PodGC) because the node it was running on no longer exists. **You can retry the anlaysis.**


# Base

## Introduction to Base

Base is a **genomics data aggregation** and **knowledge management** solution suite. It is a secure and scalable integrated genomics **data analysis solution** which provides Information management and knowledge mining. You can **analyze, aggregate and query data** for new insights that can inform and improve diagnostic assay development, clinical trials, patient testing and patient care. Clinically relevant data needs to be generated and extracted from routine clinical testing and clinical questions need to be asked across all data and information sources. As a **large data store**, Base provides a secure and compliant environment to accumulate data, allowing for efficient exploration of the aggregated data. This data consists of test results, patient data, metadata, reference data, consent and QC data.

### Use Cases

Base can be used by for different use cases:

* Clinical and Academic Researchers:
  * Big data storage solution housing all aggregated sample test outcomes
  * Analyze information by way of a convenient query formalism
  * Look for signals in combined phenotypic and genotypic data
  * Analyze QC patterns over large cohorts of patients
  * Securely share (sub)sets of data with other scientists
  * Generate reports and analyze trends in a straightforward and simple manner.
* Bioinformaticians:
  * Access, consult, audit, and query all relevant data and QC information for tests run
  * All accumulated data and accessible pipelines can be used to investigate and improve bioinformatics for clinical analysis
  * Metadata is captured via automatic pipeline version tracking, including information on individual tools and/or reference files used during processing for each sample analyzed, information on the duration of the pipeline, the execution path of the different analytical steps, or in case of failure, exit codes can be warehoused.
* Product Developers and Service Providers:
  * Better understand the efficiency of kits and tests
  * Analyze usage, understand QC data trends, improve products
  * Store and aggregate business intelligence data such as lab identification, consumption patterns and frequency, as well as allow renderings of test result outcome trends and much more.

### Base Action Possibilities

* **Data Warehouse Creation**: Build a relational database for your Project in which desired data sets can be selected and aggregated. Typical data sets include pipeline output metrics and other suitable data files generated by the Platform Core platform which can be complemented by additional public (or privately built) databases.
* **Report and Export**: Once created, a data warehouse can be mined using standard database query instructions. All Base data is stored in a structured and easily accessible way. An interface allows for the selection of specific datasets and conditional reporting. All queries can be stored, shared, and re-used in the future. This type of standard functionality supports most expected basic mining operations, such as variant frequency aggregation. All result sets can be downloaded or exported in various standard data formats for integration in other reporting or analytical applications.
* **Detect Signals and Patterns**: extensive and detailed selection of subsets of patients or samples adhering to any imaginable set of conditions is possible. Users can, for example, group and list subjects based on a combination of (several) specific genetic variants in combination with patient characteristics such as therapeutic (outcome) information. The built-in integration with public datasets allows users to retrieve all relevant publications, or clinically significant information for a single individual or a group of samples with a specific variant. Virtually any possible combination of stored sample and patient information allow for detecting signals and patterns by a simple single query on the big data set.
* **Profile/Cluster patients**: use and re-analyze patient cohort information based on specific sample or individual characteristics. For instance, they might want to run a next agile iteration of clinical trials with only patients that respond. Through integrated and structured consent information allowing for time-boxed use, combined with the capability to group subjects by the use of a simple query, patients can be stratified and combined to export all relevant individuals with their genotypic and phenotypic information to be used for further research.
* **Share your data**: Data sharing is subject to strict ethical and regulatory requirements. Base provides built-in functionality to securely share (sub)sets of your aggregated data with third parties. All data access can be monitored and audited, in this way Base data can be shared with people in and outside of an organization in a compliant and controlled fashion.

## Access

The Base module can be found at P**rojects > your\_project > Base**.\
In order to use Base, you need to meet the following requirements:

#### Subscription

Base is **included in all subscriptions**.

#### Enabling Base

Once a project is created, the project owner must navigate to **Projects > your\_project > Base** and click the **Enable** button. From that moment on, every user who has the proper permissions has access to the Base module in that project. While base creation for the project is ongoing, there will be an indication of this when trying to access **Projects > your\_project > Base > Tables**.

#### Enabling User Access

Access to the projects and Base is configured on the **Projects > your\_project > Project settings > Team** page. Here you can add or edit a user or workgroup and give them [Bench access](/project/p-team).

### Activity

The status and history of Base activities and jobs are shown on the [Activity](/project/p-activity) page.


# Tables

All tables created within Base are gathered on the **Projects > your\_project > Base > Tables** page. New tables can be created and existing tables can be updated or deleted here.

## Create a new Table

To create a new table, click **Projects > your\_project > Base > Tables > +Create**. Tables can be created from scratch or from a template that was previously saved. **Views** on data from Illumina hardware and processes are selected with the option [Import from catalogue](/project/p-base/base-tables/datacatalogue).

### Caution

{% hint style="danger" %}
**Once a table is saved it is no longer possible to edit the schema**, only new fields can be added. The workaround is switching to text mode, copying the schema of the table to which you want to make modifications and paste it into a new empty table where the necessary changes can be made before saving.
{% endhint %}

{% hint style="danger" %}
Once created, **do not try to modify your table column layout via the Query module** as even though you can execute ALTER TABLE commands, the definitions and syntax of the table will go out of sync resulting in processing issues.
{% endhint %}

### Empty Table

To create a table from scratch, complete the fields listed below and click the **Save** button. Once saved, a job will be created to create the table. To view table creation progress, navigate to the **Activity** page.

#### Table information

The table **name** is a required field and must be unique. The first character of the table must be a letter followed by letters, numbers or underscores. The **description** is optional.

#### References

Including or excluding references can be done by checking or un-checking the **Include reference** checkbox. These reference fields are not shown on the table creation page, but are added to the schema definition, which is visible after creating the table (**Projects > your\_project > Base > Tables > your\_table > Schema definition)**. By including references, additional columns will be added to the [schema](#schema) containing references to the data on the platform:

<table><thead><tr><th width="219.7734375">Reference</th><th>Originating source object reference in the Illumina platform.</th></tr></thead><tbody><tr><td><strong>data_name</strong></td><td>Original name of the data element, for example the filename.</td></tr><tr><td><strong>data_reference</strong></td><td>Reference to the data element (for example, the file id which you can see at <strong>Projects > your_project > Data > your_file > Data details</strong>).</td></tr><tr><td><strong>execution_reference</strong></td><td>reference to the pipeline execution. This field consists of the <strong>user reference</strong> in combination with the <strong>ID</strong> (uuid) of the analysis. You can see these two values on the <strong>Projects > your_project > Analysis > your_analysis > Details</strong> tab.</td></tr><tr><td><strong>analysis_id</strong></td><td>The ID (uuid) of the analysis. You can see this uuid on the <strong>Projects > your_project > Analysis > your_analysis > Details</strong> tab.</td></tr><tr><td><strong>pipeline_name</strong></td><td>Name of the pipeline. You can see this name at <strong>Projects > your_project > Flow > Pipelines > your_pipeline > Code</strong>.</td></tr><tr><td><strong>pipeline_reference</strong></td><td>(internal) Reference to the pipeline.</td></tr><tr><td><strong>pipeline_id</strong></td><td>The <strong>uuid</strong> of the pipeline, you can see this uuid at <strong>Projects > your_project > Flow > Pipelines > your_pipeline > Details > ID</strong>.</td></tr><tr><td><strong>sample_name</strong></td><td>Name of the sample.</td></tr><tr><td><strong>sample_reference</strong></td><td>(internal) Reference to the sample.</td></tr><tr><td><strong>sample_id</strong></td><td>The uuid of the sample. You can see this uuid in the <strong>URL of the sample</strong> after the /samples/ tag.</td></tr><tr><td><strong>tenant_name</strong></td><td>Name of the tenant.</td></tr><tr><td><strong>tenant_reference</strong></td><td>Reference to the tenant.</td></tr></tbody></table>

#### Schema

In an empty table, you can create a schema by adding a field with the **+Add** button for each column of the table and defining it. At any time during the creation process, it is possible to switch to the **edit definition** mode and back. The definition mode shows the JSON code, whereas the original view shows the fields in a table.

{% hint style="info" %}
**If you make a mistake in the order of columns when creating your table**, then **as long as you have not saved your table**, you can switch to **Edit definition** to change the column order. The text editor can swap or move columns whereas the built-in editor can only delete columns or add columns to the end of the sequence. When editing in text mode, it is best practice to copy the content of the text editor to a notepad before you make changes because a corrupted syntax will result in the text being wiped or reverted when switching between text and non-text mode.
{% endhint %}

Each field requires:

* a unique name (\*1) with optional description.
* a type
  * String – collection of characters
  * Bytes – raw binary data
  * Integer – whole numbers
  * Float – fractional numbers (\*2)
  * Numeric – any number (\*3)
  * Boolean – only options are “true” or “false”
  * Timestamp - Stores number of (milli)seconds passed since the Unix epoch
  * Date - Stores date in the format YYYY-MM-DD
  * Time - Stores time in the format HH:MI:SS
  * Datetime - Stores date and time information in the format YYYY-MM-DD HH:MI:SS
  * Record – has a child field
  * Variant - can store a value of any other type, including OBJECT and ARRAY
* a mode
  * Required - Mandatory field
  * Nullable - Field is allowed to have no value
  * Repeated - Multiple values are allowed in this field (will be recognized as array in Snowflake)

{% hint style="info" %}
(\*1) **Do not use reserved Snowflake keywords** such as left, right, sample, select, table,... (<https://docs.snowflake.com/en/sql-reference/reserved-keywords>) for your schema name as this will lead to SQL compilation errors.

(\*2) Float values will be exported differently depending on the output format. For example JSON will use scientific notation so verify that your consecutive processing methods support this.

(\*3) Defining the precision when creating tables with SQL is not supported as this will result in rounding issues.
{% endhint %}

### From template

Users can create their own template by making a table which is turned into a template at **Projects > your\_project > Base > Tables > your\_table > Manage (top right) > Save as template**.

If a template is created and available/active, it is possible to create a new table based on this template. The table information and references follow the rules of the empty table but in this case the schema will be pre-filled. It is possible to still edit the schema that is based on the template.

## Table information

### Table status

The status of a table can be found at **Projects > your\_project > Base > Tables**. The possible statuses are:

* **Available**: Ready to be used, both with or without data
* **Pending**: The system is still processing the table, there is probably a process running to fill the table with data
* **Deleted**: The table is deleted functionally; it still exists and can be shown in the list again by clicking the ***Show deleted tables/views*** button

Additional Considerations

* Tables created from empty data or from a template are available faster.
* When copying a table with data, it can remain in a Pending for longer periods of time.
* Clicking on the page's refresh button will update the list.

### Table details

For any available table, the following details are shown:

* **Table information**: Name, description, status, number of records and data size.

{% hint style="info" %}
The data size of tables with the same layout and content may vary slightly, depending on when and how the data was written by Snowflake.
{% endhint %}

* **Definition**: An overview of the table schema, also available in text. Fields can be added to the schema but not deleted. *Tip for deleting fields: copy the schema as text and paste in a new empty table where the schema is still editable*.
* **Preview**: A preview of the table for the 50 first rows (when data is uploaded into the table). Select **show details** to see record details.
* **Source Data**: the files that are currently uploaded into the table. You can see the Load Status of the files which can be *Prepare Started*, *Prepare Succeeded* or *Prepare Failed* and finally *Load Succeeded* or *Load Failed*.

## Table actions

From within the details of a table it is possible to perform the following actions from the Manage menu (top right) of the table:

* **Edit**: **Add fields** to the table and change the table description.
* **Copy**: Create a copy from this table in the same or a different project. In order to copy to another project, data sharing of the original project should be enabled in the details of this project. The user also has to have access to both original and target project.
* **Export as file**: Export this table as a **CSV or JSON file**. The exported file can be found in a project where the user has the access to download it.
* **Save as template**: Save the schema or an edited form of it as a template.
* **Add data**: Load additional data into the table manually. This can be done by selecting data files previously uploaded to the project, or by dragging and dropping files directly into the popup window for adding data to the table. It’s also possible to load data into a table manually or automatically via a pre-configured job. This can be done on the **Schedule** page.
* **Delete**: Delete the table.

## Manually importing data to your Table

To manually add data to your table, go to **Projects > your\_project > Base > Tables > your\_table > Manage (top right) > Add Data**

### Data selection

The data selection screen will show options to select the structure as CSV (comma-separated), TSV (tab-separated) or JSON (JavaScript Object Notation) and the location of your source data. In the first step, you select the data format and the files containing the data.

<table><thead><tr><th width="208.21875"></th><th></th></tr></thead><tbody><tr><td><strong>Data format (required)</strong></td><td>Select the format of the data which you want to import. This will also serve as filter to help you select files of this type.</td></tr><tr><td><strong>Write preference</strong></td><td>Define if data can be written to the table only when the table is empty, if the data should be appended to the table or if the table should be overwritten.</td></tr><tr><td><strong>Delimiter</strong></td><td>Which delimiter is used in the delimiter separated file. If the required delimiter is not comma, tab or pipe, select custom and define the custom delimiter.</td></tr><tr><td><strong>Custom delimiter</strong></td><td>If a custom delimiter is used in the source data, it must be defined here.</td></tr><tr><td><strong>Header rows to skip</strong></td><td>The number of consecutive header rows (at the top of the table) to skip.</td></tr><tr><td><strong>References</strong></td><td>Choose which references must be added to the table.</td></tr></tbody></table>

{% hint style="info" %}
Most of the advanced options are legacy functions and should not be used. The only exceptions are

* **Encoding**: Select if the encoding is UTF-8 (any Unicode character) or ISO-8859-1 (first 256 Unicode characters).
* **Ignore unknown values**: This applies to CSV-formatted files. You can use this function to handle **optional** fields without separators, provided that the missing fields are located at the end of the row. Otherwise, the parser can not detect the missing separator and will shift fields to the left, resulting in errors.
  * If **headers** are used: The columns that have matching fields are loaded, those that have no matching fields are loaded with NULL and remaining fields are discarded.
  * If **no headers** are used: The fields are loaded in order of occurrence and trailing missing fields are loaded with NULL, trailing additional fields are discarded.
    {% endhint %}

### Data import progress

To see the status of your data import, go to **Projects > your\_project > Activity > Base Jobs** where you will see a job of type *Prepare Data* which will have succeeded or failed. If it has failed, you can see the error message and details by double-clicking the base job. You can then take corrective actions if the input mismatched with the table design and try to run the import again (with a new copy of the file as each input file can only be used once)

If you need to cancel the import, you can do so while it is scheduled by navigating to the Base Jobs inventory and selecting the job followed by **Abort**.

### List of table data sources

To see which data has been used to populate your table go to **Projects > your\_project > Base > Tables > your\_table > Source Data**. This will list all the source data files, including those that failed to be imported. You can not use these files anymore to import again to prevent double entries. The load status will remain empty while the data is being processed and be set to load succeeded or failed after loading completes.

## How to load array data in Base

Base Table schema definitions do not include an array type, but arrays can be ingested using either the `Repeated` mode for arrays containing a single type (ie, String), or the `Variant` type.

### Parsing nested JSON data

If you have a nested JSON structure, you can import it into individual fields of your table.

```json
{
    "one": {
      "a": "1",
      "b": "1"
    },
    "three": {
      "a": "3",
      "b": "3",
      "c": "3"
    }
}
```

For example, if your JSON nested structure looks like the above and you want to get it imported into a table with a, b and c having integers as values, you need to create a matching table. This can be done either [manually](#create-a-new-table) or via the sql command `CREATE OR REPLACE TABLE json_data ( a INTEGER, b INTEGER, c INTEGER);`

Format your JSON data to have single lines per structure.

```
{"A":1,"B":1}
{"A":3,"B":3,"C":3}
```

Finally, create a [schedule](/project/p-base/base-schedule) to import your data or perform a [manual import](#manually-importing-data-to-your-table).

The resulting table will look like this:

| # | A | B | C |
| - | - | - | - |
| 1 | 1 | 1 |   |
| 2 | 3 | 3 | 3 |


# Data Catalogue

Data Catalogues provide **views** on data from Illumina hardware and processes (Instruments, Cloud software, Informatics software and Assays) so that this data can be distributed to different applications. This data consists of read-only tables to prevent updates by the applications accessing it. Access to data catalogues is included with professional and enterprise subscriptions.

## Available views

Project-level views

* ICA\_PIPELINE\_ANALYSES\_VIEW (Lists project-specific Platform Core pipeline analysis data)
* ICA\_DRAGEN\_QC\_METRIC\_ANALYSES\_VIEW (project-specific quality control metrics)

Tenant-level views

* ICA\_PIPELINE\_ANALYSES\_VIEW (Lists Platform Core pipeline analysis data)
* CLARITY\_SEQUENCINGRUN\_VIEW\_tenant (sequencing run data coming from the lab workflow software)
* CLARITY\_SAMPLE\_VIEW\_tenant (sample data coming from the lab workflow software)
* CLARITY\_LIBRARY\_VIEW\_tenant (library data coming from the lab workflow software)
* CLARITY\_EVENT\_VIEW\_tenant (event data coming from the lab workflow software)
* ICA\_DRAGEN\_QC\_METRIC\_ANALYSES\_VIEW (quality control metrics)

## Preconditions for view content

* DRAGEN metrics will only have content when DRAGEN pipelines have been executed.
* Analysis views will only have content when analyses have been executed.
* Views containing Clarity data will only have content if you have a Clarity LIMS instance with minimum version 6.0 and the Product Analytics service installed and configured. Please see the [Clarity LIMS documentation](https://support-docs.illumina.com/SW/ClarityLIMS-INT/Content/SW/ClarityLIMS/Integrations/Menu/CLPAInt.htm) for more information.
* When you use your [own AWS S3 ](/home/h-storage/s-awss3)storage in a project, metrics can not be collected and thus the DRAGEN METRICS - related views can not be used.

## Who can add or remove Catalogue data (views) to a project?

Members of a project, who have both **base contributor** and **project contributor or administrator** rights and who belong to the same tenant as the project can add views from a Catalogue. Members of a project with the same rights who do not belong to the same tenant can remove the catalogue views from a project. Therefore, if you are invited to collaborate on a project, but belong to a different tenant, you can remove catalogue views, but cannot add them again.

## Adding Catalogue data (views) to your project

To add Catalogue data,

1. Go to **Projects > your\_project > Base > Tables**.
2. Select **Add table > Import from Catalogue**.
3. A list of available views will be displayed. (Note that views which are already part of your project are not listed)
4. Select the table you want to add and choose **+Select**

Catalogue data will have View as type, the same as tables which are linked from other projects.

## Removing Catalogue data (views) from your project

To delete Catalogue data,

1. go to **Projects > your\_project > Base > Tables**.
2. Select the table you want to delete and choose **Delete**.
3. A warning will be presented to confirm your choice. Once deleted, you can add the Catalogue data again if needed.

## Catalogue table details (Catalogue Table Selection Screen)

* **View**: The name of the Catalogue table.
* **Description**: An explanation of which data is contained in the view.
* **Category**: The identification of the source system which provided the data.
* **Tenant/project**. Appended to the view name as \_tenant or \_project. Determines if the data is visible for all projects within the same tenant or only within the project. Only the tenant administrator can see the non-project views.

## Catalogue table details (Table Schema Definition)

In the **Projects > your\_project > Base > Tables** view, double-click the Catalogue table to see the details. For an overview of the available actions and details, see [Tables](/project/p-base/base-tables).

### Querying views

In this section, we provide examples of querying selected views from the Base UI, starting with **ICA\_PIPELINE\_ANALYSES\_VIEW** (project view). This table includes the following columns: TENANT\_UUID, TENANT\_ID, TENANT\_NAME, PROJECT\_UUID, PROJECT\_ID, PROJECT\_NAME, USER\_UUID, USER\_NAME, and PIPELINE\_ANALYSIS\_DATA. While the first eight columns contain straightforward data types (each holding a single value), the PIPELINE\_ANALYSIS\_DATA column is of type VARIANT, which can store multiple values in a nested structure. In SQL queries, this column returns data as a JSON object. To filter specific entries within this complex data structure, a combination of JSON functions and conditional logic in SQL queries is essential.

Since Snowflake offers robust JSON processing capabilities, the [FLATTEN](https://docs.snowflake.com/en/sql-reference/functions/flatten) function can be utilized to expand JSON arrays within the PIPELINE\_ANALYSIS\_DATA column, allowing for the filtering of entries based on specific criteria. It's important to note that each entry in the JSON array becomes a separate row once flattened. Snowflake aligns fields outside of this FLATTEN operation accordingly, i.e. the record USER\_ID in the SQL query below is "recycled".

The following query extracts

* USER\_NAME directly from the ICA\_PIPELINE\_ANALYSES\_VIEW\_project table.
* PIPELINE\_ANALYSIS\_DATA:reference and PIPELINE\_ANALYSIS\_DATA:price. These are direct accesses into the JSON object stored in the PIPELINE\_ANALYSIS\_DATA column. They extract specific values from the JSON object.
* Entries from the array 'steps' in the JSON object. The query uses LATERAL FLATTEN(input => PIPELINE\_ANALYSIS\_DATA:steps) to expand the steps array within the PIPELINE\_ANALYSIS\_DATA JSON object into individual rows. For each of these rows, it selects various elements (like bpeResourceLifeCycle, bpeResourcePresetSize, etc.) from the JSON.

Furthermore, the query filters the rows based on the status being 'FAILED' and the stepId not containing the word 'Workflow': it allows the user to find steps which failed.

```sql
SELECT
    USER_NAME as user_name,
    PIPELINE_ANALYSIS_DATA:reference as reference,
    PIPELINE_ANALYSIS_DATA:price as price,
    PIPELINE_ANALYSIS_DATA:totalDurationInSeconds as duration,
    f.value:bpeResourceLifeCycle::STRING as bpeResourceLifeCycle,
    f.value:bpeResourcePresetSize::STRING as bpeResourcePresetSize,
    f.value:bpeResourceType::STRING as bpeResourceType,
    f.value:completionTime::TIMESTAMP as completionTime,
    f.value:durationInSeconds::INT as durationInSeconds,
    f.value:price::FLOAT as price,
    f.value:pricePerSecond::FLOAT as pricePerSecond,
    f.value:startTime::TIMESTAMP as startTime,
    f.value:status::STRING as status,
    f.value:stepId::STRING as stepId
FROM
    ICA_PIPELINE_ANALYSES_VIEW_project,
    LATERAL FLATTEN(input => PIPELINE_ANALYSIS_DATA:steps) f
WHERE
    f.value:status::STRING = 'FAILED'
    AND f.value:stepId::STRING NOT LIKE '%Workflow%';
```

Now let's have a look at **DRAGEN\_METRICS\_VIEW\_project** view. Each DRAGEN pipeline on Platform Core creates multiple metrics files, e.g. SAMPLE.mapping\_metrics.csv, SAMPLE.wgs\_coverage\_metrics.csv, etc for DRAGEN WGS Germline pipeline. Each of these files is represented by a row in **DRAGEN\_METRICS\_VIEW\_project** table with columns ANALYSIS\_ID, ANALYSIS\_UUID, PIPELINE\_ID, PIPELINE\_UUID, PIPELINE\_NAME, TENANT\_ID, TENANT\_UUID, TENANT\_NAME, PROJECT\_ID, PROJECT\_UUID, PROJECT\_NAME, FOLDER, FILE\_NAME, METADATA, and ANALYSIS\_DATA. ANALYSIS\_DATA column contains the content of the file FILE\_NAME as an array of JSON objects. Similarly to the previous query we will use FLATTEN command. The following query extracts

* Sample name from the file names.
* Two metrics 'Aligned bases in genome' and 'Aligned bases' for each sample and the corresponding values.

The query looks for files SAMPLE.wgs\_coverage\_metrics.csv only and sorts based on the sample name:

```sql
SELECT DISTINCT
    SPLIT_PART(FILE_NAME, '.wgs_coverage_metrics.csv', 1) as sample_name,
    f.value:column_2::STRING as metric,
    f.value:column_3::FLOAT as value
FROM
    DRAGEN_METRICS_VIEW_project,
    LATERAL FLATTEN(input => ANALYSIS_DATA) f
WHERE
    FILE_NAME LIKE '%wgs_coverage_metrics.csv'
    AND (
        f.value:column_2::STRING = 'Aligned bases in genome'
        OR f.value:column_2::STRING = 'Aligned bases'
    )
ORDER BY
    sample_name;
```

Lastly, you can combine these views (or rather intermediate results derived from these views) using the WITH and JOIN commands. The SQL snippet below demonstrates how to join two intermediate results referred to as 'flattened\_dragen\_scrna' and 'pipeline\_table'. The query:

* Selects two metrics ('Invalid barcode read' and 'Passing cells') associated with single-cell RNA analysis from records where the FILE\_NAME ends with 'scRNA.metrics.csv', and then stores these metrics in a temporary table named 'flattened\_dragen\_scrna'.
* Retrieves metadata related to all scRNA analyses by filtering on the pipeline ID from the 'ICA\_PIPELINE\_ANALYSES\_VIEW\_project' view and stores this information in another temporary table named 'pipeline\_table'.
* Joins the two temporary tables using the JOIN operator, specifying the join condition with the ON operator.

```sql
WITH flattened_dragen_scrna AS (   
SELECT DISTINCT
    SPLIT_PART(FILE_NAME, '.scRNA.metrics.csv', 1) as sample_name,
    ANALYSIS_UUID, 
    f.value:column_2::STRING as metric,
    f.value:column_3::FLOAT as value
FROM
    DRAGEN_METRICS_VIEW_project,
    LATERAL FLATTEN(input => ANALYSIS_DATA) f
WHERE
    FILE_NAME LIKE '%scRNA.metrics.csv'
    AND (
        f.value:column_2::STRING = 'Invalid barcode read'
        OR f.value:column_2::STRING = 'Passing cells'
    )
),
pipeline_table AS (
SELECT
    PIPELINE_ANALYSIS_DATA:reference::STRING as reference,
    PIPELINE_ANALYSIS_DATA:id::STRING as analysis_id,
    PIPELINE_ANALYSIS_DATA:status::STRING as status,
    PIPELINE_ANALYSIS_DATA:pipelineId::STRING as pipeline_id,
    PIPELINE_ANALYSIS_DATA:requestTime::TIMESTAMP as start_time
FROM
    ICA_PIPELINE_ANALYSES_VIEW_project
WHERE
    PIPELINE_ANALYSIS_DATA:pipelineId = 'c9c9a2cc-3a14-4d32-b39a-1570c39ebc30'
    )
SELECT * FROM flattened_dragen_scrna JOIN pipeline_table 
ON
     flattened_dragen_scrna.ANALYSIS_UUID = pipeline_table.analysis_id;
```

### An example how to obtain the costs incurred by the individual steps of an analysis

You can use **ICA\_PIPELINE\_ANALYSES\_VIEW** to obtained the costs of individual steps of an analysis. Using the following SQL snippet you can retrieve the costs of individual steps for every analyses run in the past week.

```sql
SELECT
    USER_NAME as user_name,
    PROJECT_NAME as project,
    SUBSTRING(PIPELINE_ANALYSIS_DATA:reference, 1, 30) as reference,
    PIPELINE_ANALYSIS_DATA:status as status,
    ROUND(PIPELINE_ANALYSIS_DATA:computePrice,2) as price,
    PIPELINE_ANALYSIS_DATA:totalDurationInSeconds as duration,
    PIPELINE_ANALYSIS_DATA:startTime::TIMESTAMP as startAnalysis,
    f.value:bpeResourceLifeCycle::STRING as bpeResourceLifeCycle,
    f.value:bpeResourcePresetSize::STRING as bpeResourcePresetSize,
    f.value:bpeResourceType::STRING as bpeResourceType,
    f.value:durationInSeconds::INT as durationInSeconds,
    f.value:price::FLOAT as priceStep,
    f.value:status::STRING as status,
    f.value:stepId::STRING as stepId
FROM
    ICA_PIPELINE_ANALYSES_VIEW_project,
    LATERAL FLATTEN(input => PIPELINE_ANALYSIS_DATA:steps) f
WHERE
   PIPELINE_ANALYSIS_DATA:startTime > CURRENT_TIMESTAMP() - INTERVAL '1 WEEK'
ORDER BY
   priceStep DESC;
```

## Limitations

* Data Catalogue views cannot be shared as part of a Bundle.
* Data size is not shown for views because views are a subset of data.
* By removing Base from a project, the Data Catalogue will also be removed from that project.

## Best Practices

As tenant-level Catalogue views can contain sensitive data, it is best to save this (filtered) data to a new table and share that table instead of sharing the entire view as part of a project. To do so, add your view to a separate project and run a query on the data at **Projects > your\_project > Base > Query > New Query**. When the query completes, you can export the result as a new table. This ensures no new data will be added on consequent runs.


# Query

Queries can be used for data mining. On the **Projects > your\_project > Base > Query** page:

* New queries can be created and executed
* Already executed queries can be found in the query history
* Saved queries and query templates are listed under the saved queries tab.

## New Query

### Available tables

All available tables are listed on the **Run** tab.

{% hint style="info" %}
Metadata tables are created by syncing with the Base module. This synchronization is configured on the **Details** page within the project.
{% endhint %}

#### Input

Queries are executed using SQL (for example `Select * From table_name`). When there is a syntax issue with the query, the error will be displayed on the query screen when trying to run it. The query can be immediately executed or saved for future use.

#### Best practices and notes

{% hint style="danger" %}
**Do not use queries such as ALTER TABLE to modify your table structure as it will go out of sync with the table definition and will result in processing errors.**
{% endhint %}

* When you have duplicate column names in your query, put the columns explicitly in the select clause and use column aliases for columns with the same name.
* Case sensitive column names (such as the VARIANTS table) must be surrounded by double quotes. For example, `select * from MY_TABLE where "PROJECT_NAME" = 'MyProject'`.
* The syntax for Platform Core case-sensitive subfields is without quotes, for example `select * from MY_TABLE where ica:Tenant = 'MyTenant'` As these are case sensitive, the upper and lowercasing must be respected.
* If you want to query data from a table shared from another tenant (indicated in green), select the table (**Projects > your\_project > Base > Tables > your\_table**) to see the unique name. In the example below, the query will be `select * from demo_alpha_8298.public.TestFiles`\ <br>

  <figure><img src="/files/bCPyLVrkR1xRq7BqHkLr" alt=""><figcaption></figcaption></figure>
* For more information on queries, please also see the snowflake documentation: <https://docs.snowflake.com/en/user-guide/>

#### Querying data within columns.

Some tables contain columns with an array of values instead of a single value.

### Querying data within an array

{% hint style="info" %}
As of Platform Core version 2.27, there is a change in the use of capitals for Platform Core array fields. In previous versions, the data name within the array would start with a capital letter. As of 2.27, lowercase is used. For example `ICA:Data_reference` has become `ICA:data_reference`.

You can use the GET\_IGNORE\_CASE option to adapt existing queries when you have both data in the old syntax and new data in the lowercase syntax. The syntax is `GET_IGNORE_CASE(Table_Name.Column_Name,'Array_field')`

For example:

`select ICA:Data_reference as MY_DATA_REFERENCE from TestTable` becomes:

`select GET_IGNORE_CASE(TESTTABLE.ICA,'Data_reference') as MY_DATA_REFERENCE from TestTable`

You can also modify the data to have consistent capital usage by executing the query `update YOUR_TABLE_NAME set ica = object_delete(object_insert(ica, 'data_name', ica:Data_name), 'Data_name')` and repeating this process for all field names (Data\_name, Data\_reference, Execution\_reference, Pipeline\_name, Pipeline\_reference, Sample\_name, Sample\_reference, Tenant\_name and Tenant\_reference).
{% endhint %}

Suppose you have a table called YOUR\_TABLE\_NAME consisting of three fields. The first is a name, the second is a code and the third field is an array of data called ArrayField:

<table><thead><tr><th width="133">NameField</th><th width="127">CodeField</th><th>ArrayField</th></tr></thead><tbody><tr><td>Name A</td><td>Code A</td><td>{ “userEmail”: “email_A@server.com”, "bundleName": null, “boolean”: false }</td></tr><tr><td>Name B</td><td>Code B</td><td>{ “userEmail”: “email_B@server.com”, "bundleName": "thisbundle", “boolean”: true }</td></tr></tbody></table>

| Examples                                                                                                                                                  |
| --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| You can use the name field and code field to do queries by running                                                                                        |
| `Select * from YOUR_TABLE_NAME where NameField = "Name A"`.                                                                                               |
| If you want to **show specific data** like the email and bundle name **from the array**, this becomes                                                     |
| `Select ArrayField:userEmail as User_Email, ArrayField:bundleName as Bundle_Name from YOUR_TABLE_NAME where NameField = "Name A"`.                        |
| If you want to **use data in the array** as your selection criteria, the expression becomes                                                               |
| `Select ArrayField:userEmail as User_Email, ArrayField:bundleName as Bundle_Name from YOUR_TABLE_NAME where ArrayField:boolean = true`.                   |
| If your **criteria is text in the array**, use the `'` to delimit the text. For example:                                                                  |
| `Select ArrayField:userEmail as User_Email, ArrayField:bundleName as Bundle_Name from YOUR_TABLE_NAME where ArrayField:userEmail = 'email_A@server.com'`. |
| You can also use the **LIKE** operator with the **% wildcard** if you do not know the exact content.                                                      |
| `Select ArrayField:userEmail as User_Email, ArrayField:bundleName as Bundle_Name from YOUR_TABLE_NAME where ArrayField:userEmail LIKE '%A@server%'`       |

#### Query results

If the query is valid for execution, the result will be shown as a table underneath the input box. Only the first 200 chars of a string, record or variant field are included in the query results grid. The complete value is available through clicking the "show details" link.

From within the result page of the query, it is possible to save the result in several ways:

* **Export to > New table** saves the query result as a new table with contents.
* **Export to > New view** saves the query results as a new [view](/project/p-base/base-tables/datacatalogue#available-views).
* **Export to > Project file**: As a new table, as a view or as file to the project in CSV (Tab, Pipe or a custom delimeter is also allowed.) or JSON format. When exporting in JSON format, the result will be saved in a text file that contains a JSON object for each entry, similar to when exporting a [table](https://help.ica.illumina.com/tutorials/base_basics#export-table-data). The exported file can be located in the Data page under the folder named base\_export\_<*user\_supplied\_name*>\_<*auto generated unique id*>.

### Run a new query

1. Navigate to **Projects > your\_project > Base > Query**.
2. Enter the query to execute using SQL.
3. Select **Run**.
4. Optionally, select Save to add the query to your saved queries list.

If the **query takes more than 30 seconds** without returning a result, a message will be displayed to inform you the query is still in progress and the status can be consulted on **Projects > your\_project > Activity > Base Jobs**. Once this Query is successfully completed, the results can be found in **Projects > your\_project > Base > Query > Query History** tab.

## Query history

The query history lists all queries that were executed. Historical queries are shown with their date, executing user, returned rows and duration of the run.

1. Navigate to **Projects > your\_project > Base > Query**.
2. Select the **History** tab.
3. Select a query.
4. Perform one of the following actions:
   * **Use**—Open the query for editing and running in the Run tab. You can then select **Run** to execute the query again.
   * **Save** —Save the query to the saved queries list.
   * **View Results**—Download the results from a query or export results to a new table, view, or file in the project. Results are available for 24 hours after the query is executed. To view results after 24 hours, you need to execute the query again.

## Saved Queries

All queries saved within the project are listed under the Saved tab together with the query templates.

The saved queries can be:

* **Use —** Open the query for editing and running in the Run tab. You can then select **Run** to execute the query again.
* **Saved as template —** The saved query becomes a query template.
* **Deleted —** The query is removed from the list and cannot be opened again.

The query templates can be:

* **Opened**: This will open the query again in the “New query” tab.
* **Deleted**: The query is removed from the list and cannot be opened again.

It is possible to edit the saved queries and templates by double-clicking on each query or template. Specifically for Query Templates, the data classification can be edited to be:

* **Account**: The query template will be available for everyone within the account
* **User**: The query template will be available for the user who created it

### Run a saved Query

If you have saved a query, you can run the query again by selecting it from the list of saved queries.

1. Navigate to **Projects > your\_project > Base > Query**.
2. Select the **Saved Queries** tab.
3. Select a query.
4. Select **Open Query** to open the query in the New Query tab from where it can be edited if needed and run by selecting **Run Query**.

## Shared database for project

Shared databases are displayed under the list of Tables as Shared Database for project \<project name>.

{% hint style="info" %}
For Platform Core Cohorts customers, shared databases are available in a project Base instance. For more information on specific Cohorts shared database tables that are viewable, See [Cohorts Data in Base](/project/p-cohorts/cohorts-base).
{% endhint %}


# Schedule

On the **Schedule** page at **Projects > your\_project > Base > Schedule**, it’s possible to create a job for importing different types of data you have access to into an existing table.

When creating or editing a schedule, **Automatic import** is performed when the **Active** box is checked. **The job will run at 10 minute intervals**. In addition, for both active and inactive schedules, a **manual import** is performed when selecting the schedule and clicking the **»run** button.

## Configure a schedule

There are different types of schedules that can be set up:

* Files
* Metadata
* Administrative data.

### Files

This type will load the content of specific files from this project into a table. When adding or editing this schedule you can define the following parameters:

* **Name (required)**: The name of the scheduled job
* **Description**: Extra information about the schedule
* **File name pattern (required)**: Define in this field a part or the full name of the file name or of the tag that the files you want to upload contain. For example, if you want to import files named sample1\_reads.txt, sample2\_reads.txt, … you can fill in \_reads.txt in this field to have all files that contain \_reads.txt imported to the table.
* **Generated by Pipelines**: Only files generated by these selected pipelines are taken into account. When left clear, files from all pipelines are used.
* **Target Base Table (required)**: The table to which the information needs to be added. A drop-down list with all created tables is shown. This means the table needs to be created before the schedule can be created.
* **Write preference (required)**: Define data handling; whether it can overwrite the data
* **Data format (required)**: Select the data format of the files (CSV, TSV, JSON)
* **Delimiter (required)**: to indicate which delimiter is used in the delimiter separated file. If the delimiter is not present in list, it can be indicated as custom.
* **Active**: The job will run automatically if checked
* **Custom delimiter**: the custom delimiter that is used in the file. You can only enter a delimiter here if custom delimiter is selected.
* **Header rows to skip**: The number of consecutive header rows (at the top of the table) to skip.
* **References**: Choose which references must be added to the table
* **Advanced Options**
  * **Encoding (required)**: Select the encoding of the file.
  * **Null Marker:** Specifies a string that represents a null value in a CSV/TSV file.
  * **Quote:** The value (single character) that is used to quote data sections in a CSV/TSV file. When this character is encountered at the beginning and end of a field, it will be removed. For example, entering " as quote will remove the quotes from "bunny" and only store the word bunny itself.
  * **Ignore unknown values**: This applies to CSV-formatted files. You can use this function to handle **optional** fields without separators, provided that the missing fields are located at the end of the row. Otherwise, the parser can not detect the missing separator and will shift fields to the left, resulting in errors.
    * If **headers** are used: The columns that have matching fields are loaded, those that have no matching fields are loaded with NULL and remaining fields are discarded.
    * If **no headers** are used: The fields are loaded in order of occurrence and trailing missing fields are loaded with NULL, trailing additional fields are discarded.

### Metadata

This type will create two new tables: **BB\_PROJECT\_PIPELINE\_EXECUTIONS\_DETAIL** and **ICA\_PROJECT\_SAMPLE\_META\_DATA**. The job will load metadata (added to the samples) into **ICA\_PROJECT\_SAMPLE\_META\_DATA**. The process gathers the metadata from the samples via the data linked to the project and the metadata from the analyses in this project. Furthermore, the schedular will add provenance data to **BB\_PROJECT\_PIPELINE\_EXECUTIONS\_DETAIL**. This process gathers the execution details of all the analyses in the project: the pipeline name and status, the user reference, the input files (with identifiers), and the settings selected at runtime. This enables you to track the lineage of your data and to identify any potential sources of errors or biases. So, for example, the following query will count how many times each of the pipelines was executed and sort it accordingly:

```sql
SELECT PIPELINE_NAME, COUNT(*) AS Appearances
FROM BB_PROJECT_PIPELINE_EXECUTIONS_DETAIL
GROUP BY PIPELINE_NAME
ORDER BY Appearances DESC;
```

To obtained the similar table for the failed runs, you can execute the following SQL query:

```sql
SELECT PIPELINE_NAME, COUNT(*) AS Appearances
FROM BB_PROJECT_PIPELINE_EXECUTIONS_DETAIL
WHERE PIPELINE_STATUS = 'Failed'
GROUP BY PIPELINE_NAME
ORDER BY Appearances DESC;
```

When adding or editing this schedule you can define the following parameters:

* **Name (required)**: the name of this scheduled job.
* **Description**: Extra information about the schedule.
* **Include sensitive meta data fields**: in the meta data fields configuration, fields can be set to sensitive. When checked, those fields will also be added.
* **Active**: the job will run automatically if ticked.
* **Source (Tenant Administrators Only)**:
  * Project (default): All administrative data from this project will be added.
  * Account: All administrative data from every project in the account will be added. When a tenant admin creates the tenant-wide table with administrative data in a project and invites other users to this project, these users will see this table as well.

### Administrative data

This type will automatically create a table and load administrative data into this table. A usage overview of all executions is considered administrative data.

When adding or editing this schedule the following parameters can be defined:

* **Name (required)**: The name of this scheduled job.
* **Description**: Extra information about the schedule.
* **Include sensitive metadata fields**: In the metadata fields configuration, fields can be set to sensitive. When checked, those fields will also be added.
* **Active**: The job will run automatically if checked.
* **Source (Tenant Administrators Only)**:
  * Project (default): All administrative data from this project will be added.
  * Account: All administrative data from every project in the account will be added. When a tenant admin creates the tenant-wide table with administrative data in a project and invites other users to this project, these users will see this table as well.

### Delete schedule

Schedules can be deleted. Once deleted, they will no longer run, and they will not be shown in the list of schedules.

### Run schedule

When clicking the **Run** button, or **Save & Run** when editing, the schedule will start the job of importing the configured data in the correct tables. This way the schedule can be run manually. The result of the job can be seen in the tables. The load status is empty while the data is being processed and set to failed or succeeded once loading completes.


# Snowflake

### User

Every Base user has 1 snowflake username: ICA\_U\_\<id>

### User/Project-Bundle

For each user/project-bundle combination a role is created: ICA\_UR\_\<id>\_\<name project/bundle>\_\_\<id>

This role receives the viewer or contributor role of the project/bundle, depending on their permissions in Platform Core.

## Roles

Every project or bundle has a dedicated Snowflake database.

For each database, 2 roles are created:

* \<project/bundle name>\_\<id>\_VIEWER
* \<project/bundle name>\_\<id>\_CONTRIBUTOR

### Project viewer role

This role receives

* REFERENCE and SELECT rights on the tables/views within the project's PUBLIC schema.
* Grants on the viewer roles of the bundles linked to the project.

### Project contributor role

This role receives the following rights on current an future objects in the project's/bundle database in the PUBLIC schema:

* ownership
* select, insert, update, delete, truncate and references on tables/views/materialized views
* usage on sequences/functions/procedures/file formats
* write, read and usage on stages
* select on streams
* monitor and operate on tasks

It also receives grant on the viewer role of the project.

## Warehouses

For each project (not bundle!) 2 warehouses are created, whose size can be changed in Platform Core at **projects > your\_project > project settings > details**.

* \<projectname>\_\<id>\_QUERY
* \<projectname>\_\<id>\_LOAD

<details>

<summary>Using Load instead of Query warehouse</summary>

When you generate an oauth token, Platform Core always uses the QUERY warehouse by default (see bold part below):

*snowsql -a iap.us-east-1 -u ICA\_U\_277853 --authenticator=oauth -r ICA\_UR\_274853\_603465\_264891 -d atestbase2\_264891 -s PUBLIC **-w ATESTBASE2\_264891\_QUERY** --token=\<token>*

If you wish to use the LOAD warehouse in a session, you have 2 options :

1. Change the name in the connect string : *`snowsql -a iapdev.us-east-1 -u ICA_U_277853 --authenticator=oauth -r ICA_UR_277853_603465_264891 -d atestbase2_264891 -s PUBLIC -w ATESTBASE2_264891_`**`LOAD`**` `` ``--token=<token> `*
2. Execute the following statement after logging in : “*`use warehouse ATESTBASE2_264891_LOAD`”*

To determine which warehouse you are using, execute : `select current_warehouse()`;

</details>

### Synchronizing Tables

if you have [created tables](https://docs.snowflake.com/en/sql-reference/sql/create-table) directly in Snowflake with the OAuth token, you can synchronize them to appear in Platform Core by means of the **Projects > your\_project > Base > Tables > Sync** button.


# Bench

Platform Core provides a tool called **Bench** for interactive data analysis. This is a sandboxed workspace which runs a [docker image ](/home/h-dockerrepository)with access to the data and pipelines within a project. This workspace runs on the Amazon S3 system and comes with associated processing and provisioning costs. It is therefore best practice to **not keep your Bench instances running indefinitely**, but stopping them when not in use.

## Access

To access Bench, the following requirements must be met:

* Bench is **included** in all **subscriptions**.
* The project owner needs to [**enable Bench**](#enabling-bench-for-your-project) **for** their **project**.
* The project owner needs to give [**access** to Bench](#setting-user-level-access) to Individual **users** of that project.

### Enabling Bench for your project

After creating a project, go to **Projects > your\_project > Bench > Workspaces** page and click the **Enable** button. The entitlements you have determine the available resources for your Bench workspaces. If you have multiple entitlements, all the resources of your individual entitlements are taken into account. Once bench is enabled, users with matching [permissions](#setting-user-level-access) have access to the Bench module in that project.

{% hint style="info" %}
If you do not see the **Enable** button for Bench, then the tenant to which you belong is not the one where the project was created. Users from other tenants can create workspaces in Bench once Bench is enabled, but they cannot enable the Bench module itself.
{% endhint %}

### Setting user level access.

Once Bench has been enabled for your project, the combination of roles and teams settings determines if a user can access Bench.

* **Tenant administrators** and **project owners** are always able to access Bench and perform all actions.
* The teams settings page at **Projects > your\_project > Project Settings > Team** determines the role for the user/workgroup.
  * **No Access** means you have no access to the Bench workspace for that project.
  * **Contributor** gives you the right to start and stop the Bench workspace and to access the workspace contents, but not to create or edit the workspace.
  * **Administrator** gives you the right to create, edit, delete, start and stop the Bench workspace, and to access the actual workspace contents. In addition, the administrator can also build new derived Bench images and tools.
* Finally, a verification is done of your user rights against the required workspace permissions. You will only have access when your user rights meet or exceed the required workspace permissions. The possible required Workspace permissions include:
  * Upload / Download rights (Download rights are mandatory for technical reasons)
  * Project Level (No Access / Data Provider / Viewer / Contributor)
  * Flow (No Access / Viewer / Contributor)
  * Base (No Access / Viewer / Contributor)

<img src="/files/uZMxmaxnocnGTh3UKtKZ" alt="Flow diagram of access to Bench" width="375">


# Workspaces

The main concept in Bench is the *Workspace*. A workspace is an instance of a Docker image that runs the framework which is defined in the image (for example JupyterLab, R Studio). In this workspace, you can write and run code and graphically represent data. You can use API calls to access data, analyses, Base tables and queries in the platform. Via the command line, R-packages, tools, libraries, IGV browsers, widgets, etc. can be installed.

You can create multiple workspaces within a project and each workspace runs on an individual node and is available in different resource sizes. Each node has local storage capacity, where files and results can be temporarily stored and exported from to be permanently stored in a Project. The size of the storage capacity can range from 1GB – 16TB.

{% hint style="warning" %}
Once a workspace is started, it will be restarted every 30 days for security reasons. Even when you have automatic shutdown configured to be more than 30, the workspace will be restarted after 30 days and the remaining days will be counted in the next cycle.

You can see the remaining time until the next event (Shutdown or restart) in the workspaces overview and on the workspace details.
{% endhint %}

## Create Workspace

{% hint style="info" %}
If this is the first time you are using a workspace in a Project, click `Enable` to create new Bench Workspaces. In order to use Bench, you first need to have a workspace. This workspace determines which docker image will be used with which node and storage size.
{% endhint %}

1. Click **Projects > Your\_Project > Bench > Workspaces > + Create Workspace**
2. Complete the following fields and save the changes.

<table><thead><tr><th width="241">Field</th><th>Explanation</th></tr></thead><tbody><tr><td>Name</td><td>must be a unique name</td></tr><tr><td>Automatic Restart Reminder</td><td>The time (in days/hours) prior to an automatic restart (every 30 days) at which an email reminder for this event is sent out to the workspace owner. For example: 1d 2h</td></tr><tr><td>Automatic Shutdown</td><td>The time (in days/hours) between start of the workspace and automatic shutdown. For example: 5d 12h. When this value is more than 30 days, which is the restart period, the workspace will be restarted after 30 days and the remaining time will be counted in the next cycle. So for 50 days, you will have a restart after 30 days and then 20 days remaining before the workspace shuts down.</td></tr><tr><td>Automatic Shutdown Reminder</td><td>The time (in days/hours) prior to the automatic shutdown at which an email reminder for this event is sent out to the workspace owner. For example: 1d 2h</td></tr><tr><td>Docker image</td><td>The list of docker images includes base images from Platform Core and images uploaded to the docker repository for that domain.</td></tr><tr><td>Resource model</td><td>Size of the machine on which the workspace will run.. See <a href="/pages/rewnkXVCIALBvAXMmNdn#bench">Bench pricing</a> for available sizes.</td></tr><tr><td>Description</td><td>A place to provide additional information about the workspace.</td></tr><tr><td>Storage size</td><td>Represents the size of the storage available on the workspace. A storage <strong>from 1GB to 16TB</strong> can be provided.</td></tr><tr><td>Access<br>(available after selecting Docker image)</td><td>The options here are determined by the <a href="/pages/iqAjjCI1wfkGZqEJPjpO">Docker image settings</a>. The options you select will become available on the details tab of the Workspace when it is running. <strong>Web</strong> allows to interact with the workspace via a browser. <strong>Console</strong> provides a terminal to interact with the workspace.</td></tr><tr><td>Cluster</td><td><p>When your selected Docker image is <a href="/pages/rYcTzoM6sZEkZCfX9iLi"><strong>cluster compatible</strong></a>, the cluster settings become available. If you do not enable the cluster settings, the cluster-compatible image will be run on a single node. When enabled, you can choose if you want to use a <strong>dedicated cluster manager</strong> and if it requires web access.</p><p>You can choose to use the <strong>same Docker image for the cluster manager and cluster members</strong> (default) or to use a separate Docker image for the cluster manager and the members.</p><p>For the <strong>cluster members</strong>, you can choose to use a <strong>static</strong> amount of nodes or <strong>dynamic scaling</strong>, which resources (cpu, memory, storage and ephemeral storage) are required and if you want to run in <strong>economy mode</strong> (AWS <a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-spot-instances.html">spot instances</a>). AWS can interrupt Spot Instances with a two-minute notification which is passed on to the workspace which in turn grants a 30 second graceful shutdown period, so it is best not to use spot Instances for workloads that cannot handle individual instance interruption.<br>For more information on clusters, see <a href="/pages/rYcTzoM6sZEkZCfX9iLi">Bench Clusters</a>, <a href="/pages/o6fAZivEQX4AP7bXcf2g">Sun Grid Engine</a> / <a href="/pages/hQqiYCNeCQ5dj7hIFVmN">Spark</a>.</p></td></tr><tr><td>Internet Access</td><td>Type of access to the internet which should be provided for this workspace. <strong>Open</strong>: Internet access is allowed. <strong>Restricted</strong>: Creates a workspace with no internet access. Access to the Platform Core project data is still available in this mode.<br><strong>Whitelisted URLs</strong>: Specify URLs(*1) and paths that are allowed in a restricted workspace. Separate URLS with a new line. Only domains and subdomains in the specified URL will be allowed.</td></tr><tr><td>Permissions</td><td><p><strong>Access limited to workspace owner</strong>. (*2)(Default) When this field is selected, only the workspace owner can access the workspace. Everything created in that workspace will belong to the workspace owner. The workspace itself will have the same permission rights to the <a href="/pages/tRbg9eiA1JFbuoZeBVoH">project</a>, <a href="/pages/eF7EFBY4izdd6J5A8nyh">flow</a> and <a href="/pages/5QAEVEwNaV8OznH25ohA">base</a> as the person who starts the workspace has at that time.</p><p>If you deselect the checkbox, you can grant access to the workspace to other users. In this case, the following options become configurable:</p><ul><li><strong>Project, Flow and Base access</strong> permission. Your workspace will operate with these <a href="#workspace-permissions">permissions</a>. For security reasons, users will need to have permissions matching or exceeding what you set here, regardless of their role.</li><li><strong>Download/Upload</strong> allowed</li></ul></td></tr></tbody></table>

{% hint style="info" %}
(\*1) URLs must comply with the following rules:

* URLs can be between 1 and 263 characters including dot (`.`).
* URLs can begin with a leading dot (`.`).
* Domain and Sub-domains:
  * Can include alphanumeric characters (Letters A-Z and digits 0-9). Case insensitive.
  * Can contain hyphens (`-`) and underscores (`_`), but not as a first or last character.
  * Length between 1 and 63 characters.
* Dot (`.`) must be placed after a domain or sub-domain.
* If you use a trailing slash like in the path ftp.example.net/folder/ then you will not be able to access the path ftp.example.net/folder without the trailing slash included.
* Regex for URL : `[(http(s)?):\/\/(www\.)?a-zA-Z0-9@:%._\+~#=-]\{2,256}\.[a-z]\{2,6}\b([-a-zA-Z0-9@:%_\+.~#?&\/\/=]*)`
  {% endhint %}

{% hint style="warning" %}
(\*2) When you grant workspace access to multiple users, you need to provide an [API key](/get-started/gs-getstarted#api-keys) to access the workspace. Authenticate using `icav2 config set` command. The CLI will prompt for an `x-api-key` value. enter the API Key generated from the product dashboard. See [here](/command-line-interface/cli-authentication) for more information.
{% endhint %}

{% hint style="info" %}
(\*2) If you want to use an ODBC connection in R to base, you need to set **Access limited to workspace owner** as user credentials are only injected in the user context under the path stored in environment variable USER\_TOKEN\_PATH when this setting is enabled. Since 2.42, API keys are no longer exposed in the environment. Use the following syntax to fetch the credentials in the correct format:

`JWT_TOKEN = trimws(readLines(Sys.getenv("USER_TOKEN_PATH"), warn = FALSE)[1])` This line removes whitespaces and obtains the token file which contains the credentials.

`"Authorization" = paste0("Bearer ", JWT_TOKEN)` handles the format changes in the user credentials which are now bearer tokens as opposed to API keys in previous version.
{% endhint %}

<details>

<summary>Example URLs</summary>

The following are example URLs which will be considered valid.

example.com\
[www.example.com\\](http://www.example.com\\)
<https://www.example.com\\>
subdomain.example.com\
subdomain.example.com/folder\
subdomain.example.com/folder/subfolder\
sub-domain.example.com\
sub\_domain.example.com\
example.co.uk\
subdomain.example.co.uk\
sub-domain.example.co.uk\\

Example data science-specific whitelist compatible with restricted Bench workspaces. There are two required URLs to allow for Python pip installs:

pypi.org\
files.pythonhosted.org\
repo.anaconda.com\
conda.anaconda.org\
github.com\
cran.r-project.org\
bioconductor.org\
[www.npmjs.com\\](http://www.npmjs.com\\)
mvnrepository.com\\

</details>

The workspace can be edited on the workspace **Details** tab once it is **stopped**. The changes will be applied when the workspace is restarted.

### Workspace permissions

When **Access limited to workspace owner** is selected, only the workspace owner can access the workspace. Everything created in that workspace will belong to the workspace owner.

#### Administrator vs Contributor

* Bench **administrators** are able to **create, edit and delete workspaces** and **start** and **stop** workspaces. If their permissions match or exceed those of the workspace, they can also **access the workspace contents**.
* **Contributors** are able to **start** and **stop** workspaces and if their permissions match or exceed those of the workspace, they can also **access the workspace contents**.

<table><thead><tr><th width="126.62890625"></th><th width="119.2890625">create/edit</th><th width="91.5234375">delete</th><th width="112.13671875">start/stop</th><th>access contents</th></tr></thead><tbody><tr><td>Contributor</td><td>-</td><td>-</td><td>X</td><td>when permissions match those of the workspace</td></tr><tr><td>Administrator</td><td>X</td><td>X</td><td>X</td><td>when permissions match those of the workspace</td></tr></tbody></table>

#### Setting Workspace Permissions

The [teams](/project/p-team) setting determines if someone is an administrator or contributor, while the dedicated [permissions you set on the workspace level](#create-new-workspace) indicate what the workspace itself can and cannot do within your project. For this reason, the **users need to meet or exceed the required permissions to enter this workspace and use it**.

{% hint style="info" %}
For security reasons, the Tenant administrator and Project owner can always access the workspace.
{% endhint %}

{% hint style="info" %}
If one of your permissions is not high enough as bench contributor, you will see the following message **"You are not allowed to use this workspace as your user permissions are not sufficient compared to the permissions of this workspace"**.
{% endhint %}

The permissions that a Bench workspace can receive are the following:

* Upload rights
* Download rights (required)
* Project (No Access - Dataprovider - Viewer - Contributor)
* Flow (No Access - Viewer - Contributor)
* Base (No Access - Viewer - Contributor)

Based on these permissions, you will be able to upload or download data to your Platform Core project (upload and download rights) and will be allowed to take actions in the Project, Flow and Base modules related to the granted permission.

If you encounter issues when uploading/downloading data in a workspace, the security settings for that workspace may be set to not allow uploads and downloads. This can result in **RequestError: send request failed** and read: connection reset by peer. This is by design in restricted workspaces and thus limits data access to your project via /data/project to prevent the extraction of large amounts of (proprietary) data.

{% hint style="info" %}
Workspaces which were created before this functionality existed can be upgraded by enabling these workspace permissions. If the workspaces are not upgraded, they will continue working as before.
{% endhint %}

### Delete workspace (Bench Administrators Only)

To delete a workspace with the **GUI**, go to **Projects > your\_project > Bench > Workspaces > your\_workspace** and click “Delete”. Note that the delete option is only available when the workspace is stopped.

The workspace will not be accessible anymore, nor will it be shown in the list of workspaces. The content of it will be deleted so if there is any information that should be kept, you can either put it in a docker image which you can use to start from next time, or export it using the API.

## Start workspace

The workspace is not always accessible. It **needs to be started before it can be used**. From the moment a workspace is Running, a node with a specific capacity is assigned to this workspace. From that moment on, you can start working in your workspace.

{% hint style="info" %}
**As long as the workspace is running, the resources provided for this workspace will be charged.**
{% endhint %}

#### GUI

To start the workspace, follow the next steps:

1. Go to **Projects > your\_project > Bench > Workspaces > your\_workspace > Details**
2. Click on **Start Workspace** button
3. On the top of the details tab, the status changes to “Starting”. When you click on the **>\_Access** tab, the message “The workspace is starting” appears.
4. Wait until the status is “Running” and the “Access” tab can be opened. This can take some time because the necessary resources have to be provisioned.

You can refresh the workspace status by selecting the round refresh symbol at the top right.

Once a workspace is running, it can be manually stopped or it will be automatically shut down after the amount of time configured in the [Automatic Shutdown](#create-workspace) field. Even with automatic shutdown, it is still best practice to stop your workspace run when you no longer need it to save costs.

{% hint style="info" %}
You can **edit** running workspaces to update the **shutdown timer**, **shutdown reminder** and **auto restart** reminder.
{% endhint %}

{% hint style="info" %}
If you want to open a running workspace in a new tab, then select the link at **Projects > your\_project > Bench > Workspaces > Details tab > Access**. You can also copy the link with the copy symbol in front of the link.
{% endhint %}

## Stop workspace

**GUI**

When you exit a workspace, you can choose to stop the workspace or keep it running. Keeping the workspace running means that it will continue to use resources and **incur associated costs**. To stop the workspace, select stop in the displayed dialog. You can also stop a workspace by opening it and selecting **stop** at the top right.

**General**

Stopping the workspace will stop the notebook, but will not delete local data. Content will no longer be accessible and no actions can be performed until it is restarted. Any work that has been saved will stay stored.

{% hint style="warning" %}
Storage will continue to be charged until the workspace is deleted.\
Administrators have a delete option for the workspace in the exit screen.
{% endhint %}

The project/tenant administrator can enter and stop workspaces for their project/tenant even if they did not start those workspaces at **Projects > your\_project > Bench > Workspaces > your\_workspace > Details**. Be careful not to stop workspaces that are processing data. For security reasons, a log entry is added when a project/tenant administrator enters and exits a workspace.

You can see who is using a workspace in the workspace list view.

## Workspace Tabs

### Access tab

Once the Workspace is running, the default applications are loaded. These are defined by the start script of the docker image.

The docker images provided by Illumina will load JupyterLab by default. It also contains Tutorial notebooks that can help you get started. Opening a new terminal can be done via the Launcher, **+** button above the folder structure.

### Docker Builds tab (Bench Administrators only)

To ensure that packages (and other objects, including data) are permanently installed on a Bench image, a new Bench image needs to be created, using the BUILD option in Bench. A new image can only be derived from an existing one. The build process uses the DOCKERFILE method, where an existing image is the starting point for the new Docker Image (The FROM directive), and any new or updated packages are additive (they are added as new layers to the existing Docker file).

{% hint style="info" %}
The Dockerfile commands are all run as ROOT, so it is possible to delete or interfere with an image in such a way that the image is no longer running correctly. The image does **not** have access to any underlying parts of the platform so will not be able to harm the platform, but inoperable Bench images will have to be deleted or corrected.
{% endhint %}

In order to create a derived image, open up the image that you would like to use as the basis and select the **Build** tab.

* **Name**: By default, this is the same name as the original image and it is recommended to change the name.
* **Version**: Required field which can by any value.
* **Description**: The description for your docker image (for example, indicating which apps it contains).
* **Code**: The Docker file commands must be provided in this section.

The first 4 lines of the Docker file must NOT be edited. It is not possible to start a docker file with a different FROM directive. The main docker file commands are RUN and COPY. More information on them is available in the official Docker documentation.

Once all information is present, click the **Build** button. Note that the build process can take a while. Once building has completed, the docker image will be available on the **Data** page within the Project. If the build has failed, the log will be displayed here and the log file will be in the **Data** list.

#### Tools (Bench Administrators Only)

From within the workspace it is possible to create a tool from the Docker image.

1. Click the **Manage > Create CWL Tool** button in the top right corner of the workspace.
2. Give the tool a name.
3. Replace the description of the tool to describe what it does.
4. Add a version number for the tool.
5. Click the **Docker Build** tab.
   * Here the image that accompanies the tool will be created.
   * Change the **name** for the image.
   * Change the **version**.
   * Replace the **description** to describe what the image does.
   * Below the line where it says “#Add your commands below.” write the code necessary for running this docker image.
6. Click the **General** tab. This tab and all next tabs will look familiar from Flow. Enter the information required for the tool in each of the tabs. For more detailed instruction check out the [Tool creation ](/home/h-toolrepository#create-a-tool)section in the Flow documentation.
7. Click the **Save** button in the upper, right-hand corner to start the build process.

The building can take a while. When it has completed, the tool will be available in the Tool Repository.

#### Workspace Data

To export data from your workspace to your local machine, it is best practice to move the data in your workspace to the **/data/project/** folder so that it becomes available in your project under **projects > your\_project > Data**. Although this storage is slow, it offers read and write access and access to the content from within ICA.

* For fast read-only access, link folders with the [CLI command](/project/p-bench/bench-command-line-interface) **workspace-ctl data create-mount --mode read-only**.
* For fast read/write access, link [**non-indexed folders**](/project/p-data/non-indexed-folders) which are visible, but whose contents are not accessible from Platform Core. Use the [CLI command](/project/p-bench/bench-command-line-interface) **workspace-ctl data create-mount --mode read-write** to do so. You can not have fast read-write access to indexed folders as the indexing mechanism on those would deteriorate the performance.

Every workspace you start has a read-only **/data/.software/** folder which contains the icav2 command-line interface (and readme file).

<figure><img src="/files/ouPqFiIF4NgGKlQPYk8I" alt="" width="340"><figcaption><p>File Mapping</p></figcaption></figure>

### Activity tab

The last tab of the workspace is the activity tab. On this tab **all actions performed in the workspace are shown**. For example, the creation of the workspace, starting or stopping of the workspace,etc. The activities are shown with their date, the user that performed the action and the description of the action. This page can be used to check how long the workspace has run.

In the general Activity page of the project, there is also a *Bench activity* tab. This shows all activities performed in all workspaces within the project, even when the workspace has been deleted. The *Activity* tab in the workspace only shows the action performed in that workspace. The information shown is the same as per workspace, except that here the workspace in which the action is performed is listed as well.


# Bench Clusters

## Managing a Bench cluster

### Introduction

Workspaces can have their own dedicated **cluster** which consists of a number of nodes. First the workspace node, which is used for interacting with the cluster, is started. Once the workspace node is started, the workspace cluster can be started.

The **cluster** consists of 2 components

* The **manager node** which orchestrates the workload across the members.
* Anywhere between 0 and up to maximum 50 **member nodes**.

#### Clusters can run in two modes.

* **Static** - A static cluster has a manager node and a static number of members. At start-up of the cluster, the system ensures the **predefined number of members** are added to the cluster. These nodes will keep running as long as the entire cluster runs. The system will not automatically remove or add nodes depending on the job load. This gives the fastest resource availability, but at additional cost as unused nodes stay active, waiting for work.
* **Dynamic** - A dynamic cluster has a manager node and a dynamic number of workers up to a predefined maximum (with a hard limit of 50). Based on the job load the system will scale the number of members up or down. This saves resources as only as much worker nodes as needed to perform the work are being used.

### Configuration

You manage Bench Clusters via the Platform Core UI in **Projects > your\_project > Bench > Workspaces > your\_workspace > Details**.

The following settings can be defined for a **bench cluster:**

<table><thead><tr><th width="238.51171875">Field</th><th>Description</th></tr></thead><tbody><tr><td>Separate Docker image for cluster manager and members.</td><td>When set to true, you can select a different Docker image for the cluster manager to do the orchestration and the cluster members to do the work.</td></tr><tr><td>Docker image</td><td><em>Available when separate Docker image for cluster manager and members is selected</em>. Here, you select the Docker image of your cluster manager.</td></tr><tr><td>Web access</td><td>Enable or disable web access to the cluster manager.</td></tr><tr><td>Dedicated Cluster Manager</td><td>Use a dedicated node for the cluster manager. <strong>This reserves an entire machine (based on the selected resource model) for the cluster manager.</strong> If no dedicated cluster manager is selected, one core per cluster member is reserved for scheduling.<br>For example, if you have 2 nodes of standard-medium (4 cores) and no dedicated cluster manager, then only 6 (2×3) cores are available to run tasks, because each node reserves 1 core for cluster management.</td></tr><tr><td>Resource model</td><td><em>Available when dedicated cluster manager is selected</em>. This is the <strong>resource model on which the cluster manager</strong> will run.</td></tr><tr><td>Include ephemeral storage</td><td><em>Available when dedicated cluster manager is selected.</em> Select this to create <strong>scratch space</strong> for your nodes. Enabling it makes the <strong>storage size</strong> selector appear. Data stored in this space is deleted when the instance is terminated. When you deselect this option, the storage size is 0.</td></tr><tr><td>Storage size</td><td><em>Available when ephemeral storage is selected.</em> How much storage space (1 GB–16 TB) to reserve per node as dedicated scratch space, available at <code>/scratch</code>.</td></tr><tr><td>Docker image</td><td><em>Available when separate Docker image for cluster manager and members is selected</em>. Here, you select the Docker image of your cluster members.</td></tr><tr><td>Type</td><td>Choose between <a href="#cluster-modes">Static</a> and <a href="#cluster-modes">Dynamic</a>.</td></tr><tr><td>Number of nodes /<br>Scaling interval</td><td>For <strong>static</strong>, set the <strong>number of cluster member nodes</strong> (maximum 50). For <strong>dynamic</strong>, choose the <strong>minimum and maximum</strong> number of cluster member nodes (up to 50).</td></tr><tr><td>Resource model</td><td>The type of <a href="/pages/4Dtc3hG26b93N8JI08s3#compute-types">machine</a> on which cluster members run. For each cluster member, one machine of this type is provisioned. Consider the cost impact when running many machines with a high individual <a href="/pages/rewnkXVCIALBvAXMmNdn#compute">cost</a>.</td></tr><tr><td>Economy mode</td><td>Economy mode uses AWS <a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-spot-instances.html">Spot Instances</a>. This halves many compute iCredit rates compared to standard mode, but instances can be interrupted. See <a href="/pages/rewnkXVCIALBvAXMmNdn#compute">Pricing</a> for a list of supported resource models.</td></tr><tr><td>Include ephemeral storage</td><td>Select this to create <strong>scratch space</strong> for your nodes. Enabling it makes the <strong>storage size</strong> selector appear. Data stored in this space is deleted when the instance is terminated. When you deselect this option, the storage size is 0.</td></tr><tr><td>Storage size</td><td><em>Available when ephemeral storage is selected.</em> How much storage space (1 GB–16 TB) to reserve per node as dedicated scratch space, available at <code>/scratch</code>.</td></tr></tbody></table>

### Operations

Once the [workspace](/project/p-bench/bench-workspaces#start-workspace) is started, the cluster can be started at **Projects > your\_project > Bench > Workspaces > your\_workspace > Details** and the cluster can be stopped without stopping the workspace. Stopping the workspace will also stop all clusters in that workspace.

## Managing Data in a Bench cluster

Data in a bench workspace can be divided into three groups:

* **Workspace data** is accessible in read/write mode and can be accessed from all workspace components (workspace node, cluster manager node, cluster member nodes ) at `/data`. The size of the workspace data is defined at the creation of the workspace but can be increased when editing a workspace in the Platform Core UI. This is persistent storage and data remains when a workspace is shut down.
* **Project data** can be accessed from all workspace components at `/data/project`. Every component will have their own dedicated mount to the project. Depending on the project data permissions you will be able to access it in either Read-Only or Read-Write mode.
* **Scratch** **data** is available on the cluster members at `/scratch` and can be used to store intermediate results for a given job dedicated to that member. This is temporary storage, and all data is deleted when a cluster member is removed from the cluster.

### Fast Read-Only Access

All mounts occur in `/data/mounts/`, see [data access](/project/p-bench/bench-workspaces#workspace-data) and [workspace-ctl data](/project/p-bench/bench-command-line-interface#workspace-ctl-data).

Managing these mounts is done via the workspace cli `/data/.local/bin/workspace-ctl` in the workspace. Every node will have his dedicated mount.

For fast data access, bench offers a mount solution to expose project data on every component in the workspace. This mount provides read-only access to a given location in the project data and is optimized for high read throughput per single file with concurrent access to files. It will try to utilise the full bandwidth capacity of the node.

All mounts occur in path `/data/mounts/`

#### Show mounts

```
workspace-ctl data get-mounts
```

#### Creating a mount

For fast read-only access, link folders with the [CLI command](/project/p-bench/bench-command-line-interface) **workspace-ctl data create-mount --mode read-only**.

{% code overflow="wrap" %}

```
workspace-ctl data create-mount --mount-path /data/mounts/mydata --source /data/project/mydata
```

{% endcode %}

{% hint style="info" %}
This has the same effect as using the `--mode read-only` option because this is applied by default when using `workspace-ctl data create-mount` .
{% endhint %}

#### Removing a mount

```
workspace-ctl data delete-mount --mount-path /data/mounts/mydata
```


# Sun Grid Engine (SGE) on Platform Core Bench

## Running Jobs in a Bench SGE Cluster

Once a cluster is started, the cluster manager can be accessed from the workspace node.

### Job resources

Every cluster member has a certain capacity which is determined by the selected [Resource](/project/p-bench/bench-workspaces/bench-clusters#configuration) model for the cluster member.

The following complex values have been added to the SGE cluster environment and are requestable.

* static\_cores (default: 1)
* static\_mem (default: 2G)

These values are used to avoid oversubscription of a node which can result in Out-Of-Memory or unresponsiveness. You need to ensure these limits are not exceeded.

To ensure stability of the system, some headroom is deducted from the total node capacity.

### Scaling

These two values are used by the **SGE auto scaler** when running in **dynamic mode.** The SGE auto scaler will summarise all pending jobs and their requested resources to determine the scale up/down operation within the defined range.

Cluster members will remain in the cluster for at least 300 seconds. The Auto scaler only executes one scale up/down operation at a time and is stabilised before taking on a new operation.

{% hint style="warning" %}
Job requests that require more resources than the capacity of the selected resource model will be ignored by the auto scaler and will wait indefinitely.
{% endhint %}

The operation of the auto scaler can be monitored in the log file `/data/logs/sge-scaler.log`

### Submitting jobs

Submitting a single job

```
qsub -l static_mem=1G -l static_cores=1 /data/myscript.sh
```

Submitting a job array

```
qsub -l static_mem=1G -l static_cores=1 -t 1-100 /data/myscript.sh
```

{% hint style="info" %}
Do not limit the job concurrency amount as this will result in unused cluster members.
{% endhint %}

### Monitoring members

Listing all members of the cluster

```
qhost
```

### Managing running/pending jobs

listing all jobs in the cluster

```
qstat -f
```

Showing the details of a job.

```
qstat -f -j <jobId>
```

Deleting a job.

```
qdel <jobId>
```

### Managing executed jobs

Showing the details of an executed job.

```
qacct -j <jobId>
```

## SGE Reference documentation

SGE command line options and configuration details can be found [here](https://gridengine.eu/mangridengine/manuals.html).


# Spark on Platform Core Bench

Running a Spark application in a Bench Spark Cluster

## Running a pyspark application

The JupyterLab environment is by default configured with 3 additional kernels

* PySpark – [Local](#pyspark-local)
* PySpark – [Remote](#pyspark-remote)
* PySpark – [Remote – Dynamic](#pyspark-remote-dynamic)

When one of the above kernels is selected, the spark context is automatically initialised and can be accessed using the **sc** object.

<figure><img src="/files/RbHEWxETqQBHGHWzsH33" alt=""><figcaption></figcaption></figure>

### PySpark - Local

The PySpark - Local runtime environment launches the **spark driver locally** on the workspace node and all spark **executors** are created locally on the **same node**. It does **not require a spark cluster to run** and can be used for running smaller spark applications which don’t exceed the capacity of a single node.

The spark configuration can be found at `/data/.spark/local/conf/spark-defaults.conf`.

{% hint style="info" %}
Making changes to the configuration requires a restart of the Jupyter kernel.
{% endhint %}

### PySpark - Remote

The PySpark – Remote runtime environment launches the **spark driver locally** on the workspace node and interacts with the Manager for scheduling tasks onto **executors created across the Bench Cluster**.

This configuration will not dynamically spin up executors, hence it will **not** trigger the cluster to **auto scale** when using a Dynamic Bench cluster.

The spark configuration can be found at `/data/.spark/remote/conf/spark-defaults.conf`.

{% hint style="info" %}
Making changes to the configuration requires a restart of the Jupyter kernel.
{% endhint %}

### PySpark – Remote - Dynamic

The PySpark – Remote - Dynamic runtime environment launches the **spark driver locally** on the workspace node and interacts with the Manager for scheduling tasks onto **executors created across the Bench Cluster**.

This configuration will increase/decrease the required executors which will result into a cluster that **auto scales** using a Dynamic Bench cluster

The spark configuration can be found at `/data/.spark/remote/conf-dynamic/spark-defaults.conf`.

{% hint style="info" %}
Making changes to the configuration requires a restart of the Jupyter kernel.
{% endhint %}

## Job resources

Every cluster member has a certain capacity depending on the selection of the [Resource](/project/p-bench/bench-workspaces/bench-clusters#configuration) model for the member.

A spark application consists of 1 or more **jobs**. Each Job consists out of one or more **stages**. Each stage consists out of one or more **tasks**. Task are handled by executors and executors are run on a worker (cluster member).

The following setting define the amount of cpus needed per task

```
spark.task.cpus 1
```

The following settings define the size of a single executor which handles the execution of a task

```
spark.executor.cores 4 
spark.executor.memory 4g
```

The above example allows an executor to handle 4 tasks concurrently and share a total capacity of 4Gb of memory. Depending on the resource model chosen (e.g. standard-2xlarge) a single cluster member (worker node) is able to run multiple executors concurrently (e.g. 32 cores, 128 Gb for 8 concurrent executors on a single cluster member)

## Spark User Interface

The Spark UI can be accessed via the Cluster. The Web Access URL is displayed in the Workspace details page

This Spark UI will register all applications submitted when using one of the Remote Jupyter kernels. It will provide an overview of the registered workers (Cluster members) and the applications running in the Spark cluster.

<figure><img src="/files/U920ARLy7z5hx9CRXbaL" alt=""><figcaption></figcaption></figure>

### Spark Reference documentation

See the [apache](https://spark.apache.org/docs/3.5.6/configuration.html) website


# JupyterLab

Bench workspaces require setting a Docker image to use as the image for the workspace. Platform Core provides a default Docker image with [JupyterLab](https://jupyterlab.readthedocs.io/en/stable/) installed.

JupyterLab supports [Jupyter Notebook documents](https://ipython.org/ipython-doc/dev/notebook/notebook.html#notebook-documents) (.ipynb). Notebook documents consist of a sequence of cells which may contain executable code, markdown, headers, and raw text.

The JupyterLab Docker image contains the following environment variables:

<table><thead><tr><th width="321.58984375">Variable</th><th>Set to</th></tr></thead><tbody><tr><td>ICA_URL</td><td><code>https://ica.illumina.com/ica</code> (Platform Core server URL)</td></tr><tr><td>ICA_PROJECT_UUID</td><td>Current project UUID</td></tr><tr><td>ICA_SNOWFLAKE_ACCOUNT</td><td>Platform Core Snowflake (Base) Account ID</td></tr><tr><td>ICA_SNOWFLAKE_DATABASE</td><td>Platform Core Snowflake (Base) Database ID</td></tr><tr><td>ICA_PROJECT_TENANT_NAME</td><td>Name of the owning tenant of the project where the workspace is created.</td></tr><tr><td>ICA_STARTING_USER_TENANT_NAME</td><td>Name of the tenant of the user which last started the workspace.</td></tr><tr><td>ICA_COHORTS_URL</td><td>URL of the Cohorts web application used to support the Cohort's view</td></tr></tbody></table>

{% hint style="info" %}
To export data from your workspace to your local machine, it is best practice to move the data in your workspace to the `/data/project/` folder so that it becomes available in your project under **projects > your\_project > Data**.
{% endhint %}

## Platform Core Python Library

Included in the default JupyterLab Docker image is a python library with APIs to perform actions in Platform Core, such as add data, launch pipelines, and operate on Base tables. The python library is generated from the [Open API specification](https://ica.illumina.com/ica/api/swagger/index.html#/) using [openapi-generator](https://github.com/OpenAPITools/openapi-generator).

The Platform Core Python library API documentation can be found in folder `/etc/ica/data/ica_v2_api_docs` within the JupyterLab Docker image.

See the [Bench ICA Python Library Tutorial](/tutorials/bench-ica-python-library) for examples on using the Platform Core Python library.


# Bring Your Own Bench Image

Bench images are Docker containers tailored to run in Platform Core with the necessary permissions, configuration and resources. For more information of Docker images, please refer to <https://docs.docker.com/reference/dockerfile/>

The following steps are needed to get your bench image running in Platform Core.

![](https://documents.lucid.app/documents/6f85a665-ca2e-4cdc-825c-6d0002a6e9b8/pages/0_0?a=373\&x=36\&y=78\&w=1388\&h=129\&store=1\&accept=image%2F*\&auth=LCA%208294cf0566d7d740242d0669cedec782161771aab20b949de69921ba6c6cbc19-ts%3D1782395290)

### Requirements

You need to have Docker installed in order to build your images.

For your Docker bench image to work in Platform Core, they must run on Linux X86 architecture, have the correct user id and initialization script in the Docker file.

{% hint style="info" %}
For easy reference, you can find examples of preconfigured Bench images on the [Illumina website](https://github.com/Illumina/Bench-Example-Images) which you can copy to your local machine and edit to suit your needs.

**Bench-console** provides an example to build a minimal image compatible with Platform Core Bench to run a SSH Daemon.

**Bench-web** provides an example to build a minimal image compatible with Platform Core Bench to run a Web Daemon.

**Bench-rstudio** provides an example to build a minimal image compatible with Platform Core Bench to run a rStudio Open Source.

These examples come with information on the available parameters.
{% endhint %}

### Scripts

The following scripts must be part of your Docker bench image. Please refer to the examples from the [Illumina website](https://github.com/Illumina/Bench-Example-Images) for more details.

#### Init Script (Dockerfile)

This script copies the `ica_start.sh` file which takes care of the Initialization and termination of your workspace to the location in your project from where it can be started by Platform Core when you request to start your workspace.

{% code overflow="wrap" %}

```
# Init script invoked at start of a bench workspace
COPY --chmod=0755 --chown=root:root ${FILES_BASE}/ica_start.sh /usr/local/bin/ica_start.sh
```

{% endcode %}

#### User (Dockerfile)

The user settings must be set up so that bench runs with UID 1000.

{% code overflow="wrap" %}

```
# Bench workspaces need to run as user with uid 1000 and be part of group with gid 100
RUN adduser -H -D -s /bin/bash -h ${HOME} -u 1000 -G users ica
```

{% endcode %}

#### Shutdown Script (ica\_start.sh)

To do a clean shutdown, you can capture the sigterm which is transmitted 30 seconds before the workspace is terminated.

```
# Terminate function
function terminate() {
        # Send SIGTERM to child processes
        kill -SIGTERM $(jobs -p)

        # Send SIGTERM to waitpid
        echo "Stopping ..."
        kill -SIGTERM ${WAITPID}
}

# Catch SIGTERM signal and execute terminate function.
# A workspace will be informed 30s before forcefully being shutdown.
trap terminate SIGTERM

# Hold init process until TERM signal is received
tail -f /dev/null &
WAITPID=$!
wait $WAITPID
```

### Building a Bench Image

Once you have Docker installed and completed the configuration of your Docker files, you can build your bench image.

1. Open the command prompt on your machine.
2. Navigate to the root folder of your Docker files.
3. Execute `docker build -f Dockerfile -t mybenchimage:0.0.1 .` with mybenchimage being the name you want to give to your image and 0.0.1 replaced with the version number which you want your bench image to be. For more information on this command, *see* <https://docs.docker.com/reference/cli/docker/buildx/build/>
4. Once the image has been built, save it as docker tar file with the command `docker save mybenchimage:0.0.1 | bzip2 > ../mybenchimage-0.0.1.tar.bz2` The resulting tar file will appear next to the root folder of your docker files.

{% hint style="info" %}
*If you want to build on a mac with Apple Silicon, then the build command is docker* buildx build --platform linux/amd64 -f Dockerfile -t mybenchimage:0.0.1 .
{% endhint %}

### Upload Your Docker Image to Platform Core

1. Open Platform Core and log in.
2. Go to **Projects > your\_project > Data**.
3. For small Docker images, upload the docker image file which you generated in the previous step. For large Docker images use the [service connector](/project/p-connectivity/service-connector) to better performance and reliability to import the Docker image.
4. Select the uploaded image file and perform **Manage > Change Format**.
5. From the format list, select DOCKER and save the change.
6. Go to **System Settings > Docker Repository > Create > Image**.
7. Select the uploaded docker image and fill out the other details.
   * **Name**: The name by which your docker image will be seen in the list
   * **Version**: A version number to keep track of which version you have uploaded. In our example this was 0.0.1
   * **Description**: Provide a description explaining what your docker images does or is suited for.
   * **Type**: The type of this image is Bench. The Tool type is reserved for tool images.
   * **Cluster compatible**: Indicates if this docker images is suited for [cluster computing](/project/p-bench/bench-workspaces/bench-clusters).
   * **Access**: This setting must match the available access options of your Docker image. You can choose **web access (HTTP)**, **console access (SSH)** or both. What is selected here becomes available on the **+ New Workspace** screen. Enabling an option here which your Docker image does not support, will result in access denied errors when trying to run the workspace.
   * **Regions**: If your tenant has access to multiple regions, you can select to which regions to replicate the docker image.
8. Once the settings are entered, select **Save**. The creation of the Docker image typically takes between 5 and 30 minutes. The status of your docker image will be **partial** during creation and **available** once completed.

### Start Your Bench Image

1. Navigate to **Projects > your\_project > Bench > Workspaces**.
2. Create a new workspace with **+ Create Workspace** or edit an existing workspace.
3. Fill in the bench workspace details according to [Workspaces](/project/p-bench/bench-workspaces).
4. Save your changes.
5. Select **Start Workspace**
6. Wait for the workspace to be started and you can access it either via console or the GUI.

### Access Bench Image

Once your bench image has been started, you can access it via console, web or both, depending on your configuration.

* **Web access (HTTP)** is done from either **Projects > your\_project > Bench > Workspaces > your\_Workspace > Access tab** or from the link provided at provided in your running workspace at **Projects > your\_project > Bench > Workspaces > your\_Workspace > Details tab > Access section**.
* **Console access (SSH)** is performed from your command prompt by going to the path provided in your running workspace at **Projects > your\_project > Bench > Workspaces > your\_Workspace > Details tab > Access section**.

{% hint style="info" %}
The password needed for SSH access is any one of your personal [API keys](/get-started/gs-getstarted#api-keys)
{% endhint %}

### Command-line Interface

To execute [the commands](/project/p-bench/bench-command-line-interface), your workspace needs a way to run them such as the inclusion of an SSH daemon, be it integrated into your web access image or into your console access. There is no need to download the workspace command-line interface, you can run it from within the workspace.

### Restrictions

#### Root User

* The bench image will be instantiated as a container which will be forcedly started as user with UID 1000 and GID 100.
* You cannot elevate your permissions in a running workspace.

{% hint style="warning" %}
**Do not run containers as root** as this is bad security practice.
{% endhint %}

#### Read-only Root Filesystem

Only the following folders are writeable:

* /data
* /tmp

All other folders are mounted as read-only.

#### Network Access

For **inbound** access, the following ports on the container are publicly exposed, depending on the selection made at startup.

* Web: TCP/8888
* Console: TCP/2222

For **outbound** access, a workspace can be started in two modes:

* **Public**: Access to public IP’s is allowed using TCP protocol.
* **Restricted**: Access to list of URLs are allowed.

### Context

#### Environment Variables

At runtime, the following Bench-specific environment variables are made available to the workspace instantiated from the Bench image.

| Name                                  | Description                                                                                                                                                                    | Example Values                   |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------- |
| ICA\_WORKSPACE                        | The unique identifier related to the started workspace. This value is bound to a workspace and will never change.                                                              | 32781195                         |
| ICA\_CONSOLE\_ENABLED                 | Whether Console access is enabled for this running workspace.                                                                                                                  | true, false                      |
| ICA\_WEB\_ENABLED                     | Whether Web access is enabled for this running workspace.                                                                                                                      | true, false                      |
| USER\_TOKEN\_PATH                     | The location of the token file which contains the credentials needed for interaction with ICA. See [workspaces](/project/p-bench/bench-workspaces#create-workspace) for usage. |                                  |
| ICA\_BENCH\_URL                       | The host part of the public URL which provides access to the running workspace.                                                                                                | use1-bench.platform.illumina.com |
| ICA\_PROJECT\_UUID                    | The unique identifier related to the ICA project in which the workspace was started.                                                                                           |                                  |
| ICA\_URL                              | The ICA Endpoint URL.                                                                                                                                                          | <https://ica.illumina.com/ica>   |
| <p>HTTP\_PROXY</p><p>HTTPS\_PROXY</p> | The proxy endpoint in case the workspace was started in restricted mode.                                                                                                       |                                  |
| HOME                                  | The home folder.                                                                                                                                                               | /data                            |

#### Configuration Files

Following files and folders will be provided to the workspace and made accessible for reading at runtime.

| Name                | Description                                                                                         |
| ------------------- | --------------------------------------------------------------------------------------------------- |
| /etc/workspace-auth | Contains the SSH rsa public/private keypair which is required to be used to run the workspace SSHD. |

#### Software Files

At runtime, Platform Core-related software will automatically be made available at /data/.software in read-only mode.

New versions of Platform Core software will be made available after a restart of your workspace.

#### Important Folders

| Name            | Description                                                                                                                                                                   |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| /data           | <p>This folder contains all data specific to your workspace.</p><p>Data in this folder is not persisted in your project and will be removed at deletion of the workspace.</p> |
| /data/project   | This folder contains all your project data.                                                                                                                                   |
| /data/.software | This folder contains Platform Core-related software.                                                                                                                          |

### Bench Lifecycle

#### Workspace Lifecycle

When a bench workspace is instantiated from your selected bench image, the following script is invoked: `/usr/local/bin/ica_start.sh`

{% hint style="info" %}
This script needs to be available and executable otherwise your workspace will not boot.
{% endhint %}

This script is the main process in your running workspace and cannot run to completion as it will stop the workspace and instantiate a restart (see [init script](#init-script)).

This script can be used to invoke other scripts.

When you stop a workspace, a TERM signal is sent to the main process in your bench workspace. You can trap this signal to handle the stop gracefully (see [shutdown script)](#shutdown-script) and shut down child processes of the main process. The workspace will be forcedly shut down after 30 seconds if your main process hasn’t stopped within the given period.

### Troubleshooting

#### Build Argument

If you get the error "docker buildx build" requires exactly 1 argument when trying to build your docker image, then a possible cause is missing the last `.` of the command.

#### Server Connection Error

When you stop the workspace when users are still actively using it, they will receive a message showing a **Server Connection Error**.


# Bench Internal Command Line Interface

The following is a list of available bench CLI commands and their options. You can execute these from within a running bench workspace.

Please refer to the examples from the [Illumina website](https://github.com/Illumina/Bench-Example-Images) for more details.

## workspace-ctl

```
Usage:
  workspace-ctl [flags]
  workspace-ctl [command]

Available Commands:
  completion  Generate completion script
  compute     
  data        
  help        Help about any command
  software    
  workspace   

Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
  -h, --help               help for workspace-ctl
      --help-tree          
      --help-verbose       
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")

Use "workspace-ctl [command] --help" for more information about a command.
```

### workspace-ctl completion

```
To load completions:

Bash:

  $ source <(workspace-ctl completion bash)

  # To load completions for each session, execute once:
  # Linux:
  $ workspace-ctl completion bash > /etc/bash_completion.d/workspace-ctl
  # macOS:
  $ workspace-ctl completion bash > /usr/local/etc/bash_completion.d/workspace-ctl

Zsh:

  # If shell completion is not already enabled in your environment,
  # you will need to enable it.  You can execute the following once:

  $ echo "autoload -U compinit; compinit" >> ~/.zshrc

  # To load completions for each session, execute once:
  $ yourprogram completion zsh > "${fpath[1]}/_yourprogram"

  # You will need to start a new shell for this setup to take effect.

fish:

  $ workspace-ctl completion fish | source

  # To load completions for each session, execute once:
  $ workspace-ctl completion fish > ~/.config/fish/completions/workspace-ctl.fish

PowerShell:

  PS> workspace-ctl completion powershell | Out-String | Invoke-Expression
  # To load completions for every new session, run:
  PS> workspace-ctl completion powershell > workspace-ctl.ps1
  # and source this file from your PowerShell profile.

Usage:
  workspace-ctl completion [bash|zsh|fish|powershell]

Global Options:
      --debug          output debug logs
      --dry-run        do not send the request to server
      --help-tree      Display commands as a tree
      --help-verbose   Extended help topics and options
      --print-curl     print curl equivalent do not send the request to server
```

### workspace-ctl compute

```
Usage:
  workspace-ctl compute [flags]
  workspace-ctl compute [command]

Available Commands:
  get-cluster-details 
  get-logs            
  get-pools           
  scale-pool          

Flags:
  -h, --help           help for compute
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")

Use "workspace-ctl compute [command] --help" for more information about a command.
```

#### **workspace-ctl compute get-cluster-details**

```
Usage:
  workspace-ctl compute get-cluster-details [flags]

Flags:
  -h, --help           help for get-cluster-details
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl compute get-logs**

```
Usage:
  workspace-ctl compute get-logs [flags]

Flags:
  -h, --help           help for get-logs
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl compute get-pools**

```
Usage:
  workspace-ctl compute get-pools [flags]

Flags:
      --cluster-id string   Required. Cluster ID
  -h, --help                help for get-pools
      --help-tree           
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl compute scale-pool**

```
Usage:
  workspace-ctl compute scale-pool [flags]

Flags:
      --cluster-id string       Required. Cluster ID
  -h, --help                    help for scale-pool
      --help-tree               
      --help-verbose            
      --pool-id string          Required. Pool ID
      --pool-member-count int   Required. New pool size

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

### workspace-ctl data

```
Usage:
  workspace-ctl data [flags]
  workspace-ctl data [command]

Available Commands:
  create-mount Create a data mount under /data/mounts. Return newly created mount.
  delete-mount Delete a data mount
  get-mounts   Returns the list of data mounts

Flags:
  -h, --help           help for data
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")

Use "workspace-ctl data [command] --help" for more information about a command.
```

#### **workspace-ctl data create-mount**

```
Create a data mount under /data/mounts. Return newly created mount.

Usage:
  workspace-ctl data create-mount [flags]

Aliases:
  create-mount, mount

Flags:
  -h, --help                help for create-mount
      --help-tree           Display commands as a tree
      --help-verbose        Extended help topics and options
      --mode string         Enum:["read-only","read-write"]. Mount mode i.e. read-only, read-write
      --mount-path string   Where to mount the data, e.g. /data/mounts/hg38data (or simply hg38data)
      --source string       Required. Source data location, e.g. /data/project/myData/hg38 or fol.bc53010dec124817f6fd08da4cf3c48a (ICA folder id)
      --wait                Wait for new mount to be available on all nodes before sending response
      --wait-timeout int    Max number of seconds for wait option. Absolute max: 300 (default 300)

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl data delete-mount**

```
Delete a data mount

Usage:
  workspace-ctl data delete-mount [flags]

Aliases:
  delete-mount, unmount

Flags:
  -h, --help                help for delete-mount
      --help-tree           
      --help-verbose        
      --id string           Id of mount to remove
      --mount-path string   Path of mount to remove

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl data get-mounts**

```
Returns the list of data mounts

Usage:
  workspace-ctl data get-mounts [flags]

Aliases:
  get-mounts, list-mounts

Flags:
  -h, --help           help for get-mounts
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

### workspace-ctl help

```
Usage:
  workspace-ctl [flags]
  workspace-ctl [command]

Available Commands:
  completion  Generate completion script
  compute     
  data        
  help        Help about any command
  software    
  workspace   

Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
  -h, --help               help for workspace-ctl
      --help-tree          
      --help-verbose       
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")

Use "workspace-ctl [command] --help" for more information about a command.
```

#### **workspace-ctl help completion**

```
To load completions:

Bash:

  $ source <(yourprogram completion bash)

  # To load completions for each session, execute once:
  # Linux:
  $ yourprogram completion bash > /etc/bash_completion.d/yourprogram
  # macOS:
  $ yourprogram completion bash > /usr/local/etc/bash_completion.d/yourprogram

Zsh:

  # If shell completion is not already enabled in your environment,
  # you will need to enable it.  You can execute the following once:

  $ echo "autoload -U compinit; compinit" >> ~/.zshrc

  # To load completions for each session, execute once:
  $ yourprogram completion zsh > "${fpath[1]}/_yourprogram"

  # You will need to start a new shell for this setup to take effect.

fish:

  $ yourprogram completion fish | source

  # To load completions for each session, execute once:
  $ yourprogram completion fish > ~/.config/fish/completions/yourprogram.fish

PowerShell:

  PS> yourprogram completion powershell | Out-String | Invoke-Expression

  # To load completions for every new session, run:
  PS> yourprogram completion powershell > yourprogram.ps1
  # and source this file from your PowerShell profile.

Usage:
  workspace-ctl completion [bash|zsh|fish|powershell]

Flags:
  -h, --help   help for completion
```

#### **workspace-ctl help compute**

```
Usage:
  workspace-ctl compute [flags]
  workspace-ctl compute [command]

Available Commands:
  get-cluster-details 
  get-logs            
  get-pools           
  scale-pool          

Flags:
  -h, --help           help for compute
      --help-tree      
      --help-verbose

Use "workspace-ctl compute [command] --help" for more information about a command.
```

#### **workspace-ctl help compute get-cluster-details**

```
Usage:
  workspace-ctl compute get-cluster-details [flags]

Flags:
  -h, --help           help for get-cluster-details
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help compute get-logs**

```
Usage:
  workspace-ctl compute get-logs [flags]

Flags:
  -h, --help           help for get-logs
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help compute get-pools**

```
Usage:
  workspace-ctl compute get-pools [flags]

Flags:
      --cluster-id string   Required. Cluster ID
  -h, --help                help for get-pools
      --help-tree           
      --help-verbose
```

#### **workspace-ctl help compute scale-pool**

```
Usage:
  workspace-ctl compute scale-pool [flags]

Flags:
      --cluster-id string       Required. Cluster ID
  -h, --help                    help for scale-pool
      --help-tree               
      --help-verbose            
      --pool-id string          Required. Pool ID
      --pool-member-count int   Required. New pool size
```

#### **workspace-ctl help data**

```
Usage:
  workspace-ctl data [flags]
  workspace-ctl data [command]

Available Commands:
  create-mount Create a data mount under /data/mounts. Return newly created mount.
  delete-mount Delete a data mount
  get-mounts   Returns the list of data mounts

Flags:
  -h, --help           help for data
      --help-tree      
      --help-verbose

Use "workspace-ctl data [command] --help" for more information about a command.
```

#### **workspace-ctl help data create-mount**

```
Create a data mount under /data/mounts. Return newly created mount.

Usage:
  workspace-ctl data create-mount [flags]

Aliases:
  create-mount, mount

Flags:
  -h, --help                help for create-mount
      --help-tree           
      --help-verbose        
      --mount-path string   Where to mount the data, e.g. /data/mounts/hg38data (or simply hg38data)
      --source string       Required. Source data location, e.g. /data/project/myData/hg38 or fol.bc53010dec124817f6fd08da4cf3c48a (ICA folder id)
      --wait                Wait for new mount to be available on all nodes before sending response
      --wait-timeout int    Max number of seconds for wait option. Absolute max: 300 (default 300)
```

#### **workspace-ctl help data delete-mount**

```
Delete a data mount

Usage:
  workspace-ctl data delete-mount [flags]

Aliases:
  delete-mount, unmount

Flags:
  -h, --help                help for delete-mount
      --help-tree           
      --help-verbose        
      --id string           Id of mount to remove
      --mount-path string   Path of mount to remove
```

#### **workspace-ctl help data get-mounts**

```
Returns the list of data mounts

Usage:
  workspace-ctl data get-mounts [flags]

Aliases:
  get-mounts, list-mounts

Flags:
  -h, --help           help for get-mounts
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help help**

```
Help provides help for any command in the application.
Simply type workspace-ctl help [path to command] for full details.

Usage:
  workspace-ctl help [command] [flags]

Flags:
  -h, --help   help for help
```

#### **workspace-ctl help software**

```
Usage:
  workspace-ctl software [flags]
  workspace-ctl software [command]

Available Commands:
  get-server-metadata   
  get-software-settings 

Flags:
  -h, --help           help for software
      --help-tree      
      --help-verbose

Use "workspace-ctl software [command] --help" for more information about a command.
```

#### **workspace-ctl help software get-server-metadata**

```
Usage:
  workspace-ctl software get-server-metadata [flags]

Flags:
  -h, --help           help for get-server-metadata
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help software get-software-settings**

```
Usage:
  workspace-ctl software get-software-settings [flags]

Flags:
  -h, --help           help for get-software-settings
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help workspace**

```
Usage:
  workspace-ctl workspace [flags]
  workspace-ctl workspace [command]

Available Commands:
  get-cluster-settings   
  get-connection-details 
  get-workspace-settings 

Flags:
  -h, --help           help for workspace
      --help-tree      
      --help-verbose

Use "workspace-ctl workspace [command] --help" for more information about a command.
```

#### **workspace-ctl help workspace get-cluster-settings**

```
Usage:
  workspace-ctl workspace get-cluster-settings [flags]

Flags:
  -h, --help           help for get-cluster-settings
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help workspace get-connection-details**

```
Usage:
  workspace-ctl workspace get-connection-details [flags]

Flags:
  -h, --help           help for get-connection-details
      --help-tree      
      --help-verbose
```

#### **workspace-ctl help workspace get-workspace-settings**

```
Usage:
  workspace-ctl workspace get-workspace-settings [flags]

Flags:
  -h, --help           help for get-workspace-settings
      --help-tree      
      --help-verbose
```

### workspace-ctl software

```
Usage:
  workspace-ctl software [flags]
  workspace-ctl software [command]

Available Commands:
  get-server-metadata   
  get-software-settings 

Flags:
  -h, --help           help for software
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")

Use "workspace-ctl software [command] --help" for more information about a command.
```

#### **workspace-ctl software get-server-metadata**

```
Usage:
  workspace-ctl software get-server-metadata [flags]

Flags:
  -h, --help           help for get-server-metadata
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl software get-software-settings**

```
Usage:
  workspace-ctl software get-software-settings [flags]

Flags:
  -h, --help           help for get-software-settings
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

### workspace-ctl workspace

```
Usage:
  workspace-ctl workspace [flags]
  workspace-ctl workspace [command]

Available Commands:
  get-cluster-settings   
  get-connection-details 
  get-workspace-settings 

Flags:
  -h, --help           help for workspace
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")

Use "workspace-ctl workspace [command] --help" for more information about a command.
```

#### **workspace-ctl workspace get-cluster-settings**

```
Usage:
  workspace-ctl workspace get-cluster-settings [flags]

Flags:
  -h, --help           help for get-cluster-settings
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl workspace get-connection-details**

```
Usage:
  workspace-ctl workspace get-connection-details [flags]

Flags:
  -h, --help           help for get-connection-details
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```

#### **workspace-ctl workspace get-workspace-settings**

```
Usage:
  workspace-ctl workspace get-workspace-settings [flags]

Flags:
  -h, --help           help for get-workspace-settings
      --help-tree      
      --help-verbose

Global Flags:
      --X-API-Key string   
      --base-path string   For example: / (default "/")
      --config string      config file path
      --debug              output debug logs
      --dry-run            do not send the request to server
      --hostname string    hostname of the service (default "api:8080")
      --print-curl         print curl equivalent do not send the request to server
      --scheme string      Choose from: [http] (default "http")
```


# Pipeline Development in Bench (Experimental)

## Introduction

The **Pipeline Development Kit** in Bench makes it easy to create Nextflow pipelines for Platform Core Flow. This kit consists of a number of development tools which are installed in `/data/.software` (regardless of which Bench image is selected) and provides the following features:

* Import to Bench
  * From public nf-core pipelines
  * From existing Platform Core Flow Nextflow pipelines
* Run in Bench
* Modify and re-run in Bench, providing fast development iterations
* Deploy to Flow
* Launch validation in Flow

## Prerequisites

* Recommended **workspace size**: Nf-core Nextflow pipelines typically require **4** or more **cores** to run.
* The pipeline development tools require
  * **Conda** which is automatically installed by “pipeline-dev” if `conda-miniconda.installer.ica-userspace.sh` is present in the image.
  * **Nextflow** (version 24.10.2 is automatically installed using conda, or you can use other versions)
  * **git** (automatically installed using conda)
  * **jq, curl** (which should be made available in the image)

## NextFlow Requirements / Best Practices

Pipeline development tools work best when the following items are defined:

* Nextflow profiles:
  * ***test*** profile, specifying inputs appropriate for a validation run
  * ***docker*** profile, instructing NextFlow to use Docker
* **nextflow\_schema.json**, as described [here](https://nf-co.re/docs/nf-core-tools/pipelines/schema). This is useful for the launch UI generation. The nf-core CLI tool (installable via `pip install nf-core`) offers extensive help to create and maintain this schema.

Platform Core Flow adds one additional constraint. The **output directory** `out` is the only one automatically copied to the Project data when a Platform Core Flow analysis completes. The **`-outdir`** parameter recommended by nf-core should therefore be set to`--outdir=out` when running as a Flow pipeline.

## Pipeline Development Tools

{% hint style="info" %}
New Bench pipeline development tools only become active after a workspace reboot.
{% endhint %}

These are installed in `/data/.software` (which should be in your `$PATH`), the `pipeline-dev` script is the front-end to the other `pipeline-dev-*` tools.

Pipeline-dev fulfils a number of roles:

* Checks that the environment contains the **required tools** (conda, nextflow, etc) and offers to install them if needed.
* Checks that the **fast data mounts are present** (/data/mounts/project etc.) – it is useful to check regularly, as they get unmounted when a workspace is stopped and restarted.
* Redirects **stdout and stderr** to `.pipeline-dev.log`, with the history of log files kept as `.pipeline-dev.log.<log date>`.
* Launches the appropriate **sub-tool**.
* **Prints out errors** with backtrace, to help report issues.

***

## Usage

### 1) Starting a new Project

A pipeline-dev project relies on the following **Folder structure**, which is auto-generated when using the `pipeline-dev import*` tools.

{% hint style="warning" %}
If you start a project manually, you must follow the same folder structure.
{% endhint %}

* **Project base folder**
  * **nextflow-src**: Platform-agnostic Nextflow code, for example the github contents of an nf-core pipeline, or your usual nextflow source code.
    * **main.nf**
    * **nextflow\.config**
    * **nextflow\_schema.json**
  * **pipeline-dev.project-info**: contains project name, description, etc.
  * **nextflow-bench.config** (automatically generated when needed): contains definitions for bench.
  * **ica-flow-config**: Directory of files used when deploying pipeline to Flow.
    * **inputForm.json** (if not present, gets generated from nextflow-src/nextflow\_schema.json): input form as defined in ICA Flow.
    * **onSubmit.js**, **onRender.js** (optional, generated at the same time as inputForm.json): javascript code to go with the input form.
    * **launchPayload\_inputFormValues.json** (if not present, gets generated from the test profile): used by “pipeline-dev launch-validation-in-flow”.

#### Pipeline Sources

{% tabs %}
{% tab title="Starting from Scratch" %}
The above-mentioned project structure must be generated manually. The nf-core CLI tools can assist to generate the `nextflow_schema.json`. Tutorial [Pipeline from Scratch](/project/p-bench/pipeline-development-in-bench-experimental/creating-a-pipeline-from-scratch) goes into more details about this use case.
{% endtab %}

{% tab title="Importing nf-core Pipeline" %}

```
$ pipeline-dev import-from-nextflow <repo name e.g. nf-core/demo>
```

A directory with the same name as the nextflow/nf-core pipeline is created, and the Nextflow files are pulled into the `nextflow-src` subdirectory.

Tutorial [Nf Core Pipelines](/project/p-bench/pipeline-development-in-bench-experimental/nf-core-pipelines) goes into more details about this use case.
{% endtab %}

{% tab title="Importing an Existing Platform Core Pipeline" %}

```
$ pipeline-dev import-from-flow [--analysis-id=…] 
```

A directory called `imported-flow-analysis` is created and the analysis+pipeline assets are downloaded.

Tutorial [Updating an Existing Flow Pipeline](/project/p-bench/pipeline-development-in-bench-experimental/updating-an-existing-flow-pipeline) goes into more details about this use case.

{% hint style="info" %}
Currently only pipelines with publicly available Docker images are supported. Pipelines with Platform Core-stored images are not yet supported.
{% endhint %}
{% endtab %}
{% endtabs %}

***

### 2) Running in Bench

```
$ pipeline-dev run-in-bench [--local|--sge] 
```

Optional parameters `--local / --sge` can be added to force the execution on the local workspace node, or on the workspace cluster (when available). Otherwise, the presence of a cluster is automatically detected and used.

The script then launches nextflow. The **full nextflow command line is printed and launched**.

In case of errors, full logs are saved as `.pipeline-dev.log`

{% hint style="info" %}
Currently, not all corner cases are covered by command line options. Please start from the nextflow command printed by the tool and extend it based on your specific needs.
{% endhint %}

#### Output Example

<figure><img src="/files/8Wba0SYd0B3LNaboitMm" alt=""><figcaption><p>Nextflow output</p></figcaption></figure>

#### Container (Docker) images

Nextflow can run processes with and without Docker images. In the context of pipeline development, the pipeline-dev tools assume Docker images are used, in particular during execution with the `nextflow --profile docker`.

In NextFlow, Docker images can be **specified at the&#x20;*****process*****&#x20;level**

* This is done with the `container "<image_name:version>"` directive, which can be specified.
  * in **nextflow config** files (preferred method when following the nf-core best practices).
  * or at the start of each process definition.
* Each process can use a different docker image.
* **It is highly recommended to always specify an image.** If no Docker image is specified, Nextflow will report this. In Platform Core, a basic image will be used but with no guarantee that the necessary tools are available.

Resources such as #cpu and memory can be specified as described [here](https://nf-co.re/docs/usage/getting_started/configuration#max-resources) See [containers](https://www.nextflow.io/docs/latest/container.html) or our [tutorials](#tutorials) for details about Nextflow-Docker syntax.

Bench can push/pull/create/modify Docker images, as described in [Containers](/project/p-bench/containers-in-bench).

***

### 3) Deploying to Platform Core Flow

```
$ pipeline-dev deploy-as-flow-pipeline [--create|--update] 
```

This command does the following:

1. Generate the JSON file describing the **Platform Core Flow user interface**.
   * If *`ica-flow-config/inputForm.json`* doesn’t exist: generate it from *`nextflow-src/nextflow_swagger.json` .*
2. Generate the JSON file containing the **validation launch inputs**.
   * If *`ica-flow-config/launchPayload_inputFormValues.json`* doesn’t exist: generate it from `nextflow --profile test` inputs.
   * If **local files** are used as validation inputs or as default input values:
     * copy them to `/data/project/pipeline-dev-files/temp` .
     * get their Platform Core file IDs.
     * use these file ids in the launch specifications.
   * If **remote files** are used as validation inputs or as default input values of an input of type “file” (and not “string”): do the same as above.
3. **Identify the pipeline name** to use for this new pipeline deployment:
   * If a deployment has already occurred in this project, or if the project was imported from an existing Flow pipeline, start from this pipeline name. Otherwise start from the project name.
   * Identify which already-deployed pipelines have the same base name, with or without suffixes that could be some versioning (\_v\<number>, \_\<number>, \_\<date>).
   * Ask the user if they prefer to update the current version of the pipeline, create a new version, or enter a new name of their choice – or use the `--create/--update` parameters when specified, for scripting without user interactions.
4. A new **Platform Core Flow pipeline gets created** (except in case of pipeline update) .
   * The current Nextflow version in Bench is used to select the best Nextflow version to be used in Flow.
5. `nextflow-src` **folder is uploaded** file by file as pipeline assets.

Output Example:

<figure><img src="/files/CCD1eULIc60BOgK8Mqc5" alt=""><figcaption></figcaption></figure>

The pipeline name, id and URL are printed out, and if your environment allows, Ctrl+Click/Option+Click/Right click can open the URL in a browser.

Opening the URL of the pipeline and clicking on **Start Analysis** shows the generated user interface:

<figure><img src="/files/WilF41GRg7Z15tUmedWx" alt=""><figcaption></figcaption></figure>

***

### 4) Launching Validation in Flow

```
$ pipeline-dev launch-validation-in-flow 
```

The *`ica-flow-config/launchPayload_inputFormValues.json`* file generated in the previous step is submitted to Platform Core Flow to **start an analysis** with the same validation inputs as “nextflow --profile test”.

Output Example:

<figure><img src="/files/BGNGpfsqLBRMMGo0yl6t" alt=""><figcaption><p>launch-validation-in-flow</p></figcaption></figure>

The analysis name, id and URL are printed out, and if your environment allows, Ctrl+Click/Option+Click/Right click can open the URL in a browser.

***

## Tutorials

* [Creating a Pipeline from Scratch](/project/p-bench/pipeline-development-in-bench-experimental/creating-a-pipeline-from-scratch)
* [nf-core Pipelines](/project/p-bench/pipeline-development-in-bench-experimental/nf-core-pipelines)
* [Updating an Existing Flow Pipeline](/project/p-bench/pipeline-development-in-bench-experimental/updating-an-existing-flow-pipeline)


# Creating a Pipeline from Scratch

## Introduction

This tutorial shows you how to start a new pipeline from scratch

* [prepare linux tool + validation inputs](#preparation)
* [wrap in Nextflow](#wrapping-in-nextflow)
* [wrap the pipeline in Bench](#wrap-the-pipeline-in-bench)
* [deploy pipeline as an ICA Flow pipeline](#deploy-as-a-flow-pipeline)
* [launch Flow validation test from Bench](#run-validation-test-in-flow)

***

## Preparation

Start Bench workspace

* For this tutorial, any **instance size** will work, even the smallest *standard-small.*
* Select the **single user workspace** permissions (aka "Access limited to workspace owner "), which allows us to deploy pipelines.
* A small amount of **disk space** (10GB) will be enough.

We are going to wrap the "gzip" linux compression tool with inputs:

* 1 file
* compression level: integer between 1 and 9

{% hint style="info" %}
We intentionally do not include sanity checks, to keep this scenario simple.
{% endhint %}

#### Creation of test file:

```
mkdir demo_gzip
cd demo_gzip
echo test > test_input.txt
```

***

## Wrapping in Nextflow

Here is an example of NextFlow code that wraps the bzip2 command and publishes the final output in the “out” folder:

```
mkdir nextflow-src
# Create nextflow-src/main.nf using contents below
vi nextflow-src/main.nf
```

#### nextflow-src/main.nf

```
nextflow.enable.dsl=2
 
process COMPRESS {
  input:
    path input_file
    val compression_level
 
  output:
    path "${input_file.simpleName}.gz" // .simpleName keeps just the filename
    publishDir 'out', mode: 'symlink'
 
  script:
    """
    gzip -c -${compression_level} ${input_file} > ${input_file.simpleName}.gz
    """
}
 
workflow {
    input_path = file(params.input_file)
    gzip_out = COMPRESS(input_path, params.compression_level)
}
```

Save this file as nextflow-src/main.nf, and check that it works:

```
nextflow run nextflow-src/ --input_file test_input.txt --compression_level 5
```

#### Result

<figure><img src="/files/POMB9giPgR7dtTBiBqCt" alt=""><figcaption></figcaption></figure>

***

## Wrap the Pipeline in Bench

We now need to:

* Use Docker
* Follow some nf-core best practices to make our source+test compatible with the pipeline-dev tools

### **Using Docker:**

In NextFlow, Docker images can be specified at the *process* level

* Each **process** may use a **different docker image**
* It is highly recommended to **always specify an image**. If no Docker image is specified, Nextflow will report this. In ICA, a basic image will be used but with no guarantee that the necessary tools are available.

**Specifying the Docker image** is done with the `container '<image_name:version>'` directive, which can be specified

* at the start of each process definition
* or in nextflow config files (preferred when following nf-core guidelines)

For example, create ***nextflow-src/nextflow\.config***:

```
process.container = 'ubuntu:latest'
```

We can now run with nextflow's `-with-docker` option:

{% code overflow="wrap" %}

```
nextflow run nextflow-src/ --input_file test_input.txt --compression_level 5 -with-docker
```

{% endcode %}

Following some nf-core [best practices ](https://nf-co.re/docs/usage/getting_started/configuration)to make our source+test compatible with the pipeline-dev tools:

### Create NextFlow “test” profile <a href="#id-202503benchasdevenvnewpipelinefromscratch-wrappinglinuxtool-createnextflow-test-profile" id="id-202503benchasdevenvnewpipelinefromscratch-wrappinglinuxtool-createnextflow-test-profile"></a>

Here is an example of “test” profile that can be added to `nextflow-src/nextflow.config` to define some input values appropriate for a validation run:

#### nextflow-src/nextflow\.config

```
process.container = 'ubuntu:latest'
 
profiles {
  test {
    params {
      input_file = 'test_input.txt'
      compression_level = 5
    }
  }
}
```

With this profile defined, we can now run the same test as before with this command:

```
nextflow run nextflow-src/ -profile test -with-docker
```

### Create NextFlow “docker” profile <a href="#id-202503benchasdevenvnewpipelinefromscratch-wrappinglinuxtool-createnextflow-docker-profile" id="id-202503benchasdevenvnewpipelinefromscratch-wrappinglinuxtool-createnextflow-docker-profile"></a>

A “docker” profile is also present in all nf-core pipelines. Our pipeline-dev tools will make use of it, so let’s define it:

#### nextflow-src/nextflow\.config

```
process.container = 'ubuntu:latest'
 
profiles {
  test {
    params {
      input_file = 'test_input.txt'
      compression_level = 5
    }
  }
 
  docker {
    docker.enabled = true
  }
}
```

We can now run the same test as before with this command:

```
nextflow run nextflow-src/ -profile test,docker
```

We also have enough structure in place to start using the pipeline-dev command:

```
pipeline-dev run-in-bench
```

In order to deploy our pipeline to ICA, we need to generate the user interface input form.

This is done by using nf-core's recommended `nextflow_schema.json.`

For our simple example, we generate a minimal one by hand (done by using one of the nf-core pipelines as example):

#### nextflow-src/nextflow\_schema.json

```
{
    "$defs": {
        "input_output_options": {
            "title": "Input/output options",
            "properties": {
                "input_file": {
                    "description": "Input file to compress",
                    "help_text": "The file that will get compressed",
                    "type": "string",
                    "format": "file-path"
                },
                "compression_level": {
                    "type": "integer",
                    "description": "Compression level to use (1-9)",
                    "default": 5,
                    "minimum": 1,
                    "maximum": 9
               }
            }
        }
    }
}
```

In the next step, this gets converted to the `ica-flow-config/inputForm.json` file.

{% hint style="info" %}
Note: For large pipelines, as described on the nf-core [website](https://nf-co.re/docs/nf-core-tools/pipelines/schema)

> Manually building JSONSchema documents is not trivial and can be very error prone. Instead, the nf-core pipelines schema build command collects your pipeline parameters and gives interactive prompts about any missing or unexpected params. If no existing schema is found it will create one for you.

We recommend looking into "*nf-core pipelines schema build -d nextflow-src/"*, which comes with a web builder to add descriptions etc.
{% endhint %}

***

## Deploy as a Flow Pipeline

We just need to create a final file, which we had skipped until now: Our **project description file**, which can be created via the command `pipeline-dev project-info --init`:

#### pipeline-dev.project\_info

```
$ pipeline-dev project-info --init
 
pipeline-dev.project-info not found. Let's create it with 2 questions:
 
Please enter your project name: demo_gzip
Please enter a project description: Bench gzip demo
```

We can now run:

```
pipeline-dev deploy-as-flow-pipeline
```

After generating the ICA-Flow-specific files in the `ica-flow-config` folder (JSON input specs for Flow launch UI + list of inputs for next step's validation launch), the tool identifies which previous versions of the same pipeline have already been deployed (in ICA Flow, pipeline versioning is done by including the version number in the pipeline name).

It then asks if we want to update the latest version or create a new one.

**Choose "3"** and enter a name of your choice to avoid conflicts with all the others users following this same tutorial.

At the end, the URL of the pipeline is displayed. If you are using a terminal that supports it, Ctrl+click or middle-click can open this URL in your browser.

<figure><img src="/files/lH0o4JnOU8zlEi1u500Z" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/Q2brXOqAOTKWsmgqTTJ6" alt=""><figcaption></figcaption></figure>

***

## Run Validation Test in Flow

```
pipeline-dev launch-validation-in-flow
```

This **launches** **an** **analysis** in ICA Flow, using the same inputs as the pipeline's "test" profile.

Some of the input files will have been copied to your ICA project in order for the analysis launch to work. They are stored in the folder `/data/project/bench-pipeline-dev/temp-data`.

#### Result

```
/data/demo $ pipeline-dev launch-validation-in-flow

pipelineld: 331f209d-2a72-48cd-aa69-070142f57f73
Getting Analysis Storage Id
Launching as ICA Flow Analysis...
ICA Analysis created:
- Name: Test demo_gzip
- Id: 17106efc-7884-4121-a66d-b551a782b620
- Url: https://stage.v2.stratus.illumina.com/ica/projects/1873043/analyses/17106efc-7884-4121-a66d-b551a782620
```

<figure><img src="/files/8FNzEFzGQHA7hMGcevQy" alt=""><figcaption></figcaption></figure>


# nf-core Pipelines

## Introduction

This tutorial shows you how to

* [Import any nf-core pipeline from their public repository.](#import-nf-core-pipeline-to-bench)
* [Run the pipeline in Bench.](#run-validation-test-in-bench)
  * [monitor](#monitoring) the execution
* [Deploy pipeline as an ICA Flow pipeline](#deploy-as-flow-pipeline).
* [Launch Flow validation test from Bench.](#run-validation-test-in-flow)

## Preparation

* Start Bench workspace
  * For this tutorial, the **instance size** depends on the flow you import, and whether you use a Bench cluster:
    * If using a **cluster**, choose *standard-small or standard-medium* for the workspace master node
    * **Otherwise**, choose at least *standard-large* as nf-core pipelines often need more than 4 cores to run.
  * Select the **single user workspace** permissions (aka "Access limited to workspace owner "), which allows us to deploy pipelines
  * Specify at least 100GB of **disk space**
* Optional: After choosing the image, **enable a cluster** with at least this one *standard-large*instance type
* **Start the workspace**, then (if applicable) start the cluster

## Import nf-core Pipeline to Bench

```
mkdir demo
cd demo
pipeline-dev import-from-nextflow nf-core/demo
```

If conda and/or nextflow are not installed, *pipeline-dev* will offer to install them.

The Nextflow files are pulled into the `nextflow-src` subfolder.

{% hint style="info" %}
A larger example that still runs quickly is *nf-core/sarek*
{% endhint %}

#### Result

```
/data/demo $ pipeline-dev import-from-nextflow nf-core/demo

Creating output folder nf-core/demo
Fetching project nf-core/demo

Fetching project info
project name: nf-core/demo
repository  : https://github.com/nf-core/demo
local path  : /data/.nextflow/assets/nf-core/demo
main script : main.nf
description : An nf-core demo pipeline
author      : Christopher Hakkaart

Pipeline “nf-core/demo” successfully imported into nf-core/demo.

Suggested actions:
  cd nf-core/demo
  pipeline-dev run-in-bench
  [ Iterative dev: Make code changes + re-validate with previous command ]
  pipeline-dev deploy-as-flow-pipeline
  pipeline-dev launch-validation-in-flow
```

## Run Validation Test in Bench

All nf-core pipelines conveniently define a "test" profile that specifies a set of validation inputs for the pipeline.

The following command **runs this test profile**. If a Bench cluster is active, it runs on your Bench cluster, otherwise it runs on the main workspace instance.

```
cd nf-core/demo
pipeline-dev run-in-bench
```

{% hint style="info" %}
The pipeline-dev tool is using "nextflow run ..." to run the pipeline. The full nextflow command is printed on stdout and can be copy-pasted+adjusted if you need additional options.
{% endhint %}

#### Result

<figure><img src="/files/SWqJZNdqznBX6608FkuJ" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/eDGL2N7Paix3Ed1KQbFV" alt=""><figcaption></figcaption></figure>

#### Monitoring

When a pipeline is running **locally** (i.e. not on a Bench cluster), you can monitor the task execution from another terminal with `docker ps`

When a pipeline is running on **your Bench cluster**, a few commands help to monitor the tasks and cluster. In another terminal, you can use:

* `qstat` to see the tasks being pending or running
* **`tail /data/logs/sge-scaler.log.`**`<latest available workspace reboot time>` to check if the cluster is scaling up or down (it currently takes 3 to 5 minutes to get a new node)

<figure><img src="/files/fV6nCm7ZJNUO7vjdqYg5" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/IrWl3GzQQKSowOVr4tRg" alt=""><figcaption></figcaption></figure>

#### Data Locations

* The output of the pipeline is in the `outdir` folder
* Nextflow work files are under the `work` folder
* Log files are `.nextflow.log*` and `output.log`

<figure><img src="/files/Kqp9wAwiwlmfniWu2kDp" alt="" width="375"><figcaption></figcaption></figure>

## Deploy as Flow Pipeline

```
pipeline-dev deploy-as-flow-pipeline
```

After generating a few ICA-specific files (JSON input specs for Flow launch UI + list of inputs for next step's validation launch), the tool identifies which previous versions of the same pipeline have already been deployed (in ICA Flow, pipeline versioning is done by including the version number in the pipeline name, so that's what is checked here). It then asks if you want to update the latest version or create a new one.

Choose "3" and enter a name of your choice to avoid conflicts with other users following this same tutorial.

```
Choice: 3
Creating ICA Flow pipeline dev-nf-core-demo_v4
Sending inputForm.json
Sending onRender.js
Sending main.nf
Sending nextflow.config
```

At the end, the URL of the pipeline is displayed. If you are using a terminal that supports it, Ctrl+click or middle-click can open this URL in your browser.

<figure><img src="/files/Jl0JXFB1W72r5OotgHVY" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/MzHFijyIgK7NKEdzgv0M" alt=""><figcaption></figcaption></figure>

## Run Validation Test in Flow

```
pipeline-dev launch-validation-in-flow
```

This launches an analysis in ICA Flow, using the same inputs as the nf-core pipeline's "test" profile.

Some of the input files will have been copied to your ICA project to allow the launch to take place. They are stored in the folder `bench-pipeline-dev/temp-data.`

<figure><img src="/files/5gfXO7DKugHqizzn4gsA" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/MgD6DtdtLPWJYBNjkkov" alt=""><figcaption></figcaption></figure>

## Hints

<details>

<summary>Using older versions of Nextflow</summary>

Some older nf-core flows are still using DSL1, which is only working up to Nextflow 22.

An easy solution is to create a conda environment for nextflow 22:

```
conda create -n nextflow22
 
# If, like me, you never ran "conda init", do it now:
conda init
bash -l # To load the conda's bashrc changes
 
conda activate nextflow22
conda install -y nextflow=22
 
# Check
nextflow -version
 
# Then use the pipeline-dev tools as in the demo
```

</details>


# Updating an Existing Flow Pipeline

## Introduction

This tutorial shows you how to

* [import an existing ICA Flow pipeline](#import-existing-pipeline-and-analysis-pipeline-to-bench) with a supporting validation analysis
* [run the pipeline in Bench](#run-validation-test-in-bench)
  * [monitor](#monitring) the execution
* Iterative development: [modify pipeline code ](#modify-pipeline)and validate in Bench
  * Modify [nextflow](#modify-pipeline) code
  * Modify Docker image contents ([Dockerfile](#docker-image-update-dockerfile-method) or [Interactive](#docker-image-update-interactive-method) method)
* [redeploy pipeline to ICA Flow](#deploy-as-flow-pipeline)
* [launch Flow validation test from Bench](#run-validation-test-in-flow)

## Preparation

Make sure you have access in ICA Flow to:

* the pipeline you want to work with
* an analysis exercising this pipeline, preferably with a short execution time, to use as validation test

## Start Bench Workspace

For this tutorial, the instance size depends on the flow you import, and whether you use a Bench cluster:

* When using a **cluster**, choose *standard-small or standard-medium* for the workspace master node
* **Otherwise**, choose at least *standard-large* if you re-import a pipeline that originally came from nf-core, as they typically need **4 or more CPUs** to run.
* Select the "**single user workspace**" permissions (aka "Access limited to workspace owner "), which allows us to deploy pipelines
* Specify at least **100GB of disk space**
* Optional: After choosing the image, enable a cluster with at least one *standard-large* instance type.
* **Start the workspace**, then (if applicable) also start the cluster

## Import Existing Pipeline and Analysis to Bench

```
mkdir demo-flow-dev
cd demo-flow-dev
 
pipeline-dev import-from-flow
 or
pipeline-dev import-from-flow --analysis-id=9415d7ff-1757-4e74-97d1-86b47b29fb8f
```

The starting point is the analysis id that is used as pipeline validation test (the pipeline id is obtained from the analysis metadata).

If no --analysis-id is provided, the tool lists all the successfull analyses in the current project and lets the developer pick one.

* If conda and/or nextflow are not installed, *pipeline-dev* will offer to install them.
* A folder called `imported-flow-analysis` is created.
* Pipeline Nextflow assets are downloaded into the `nextflow-src` sub-folder.
* Pipeline input form and associated javascript are downloaded into the `ica-flow-config` sub-folder.
* Analysis input specs are downloaded to the `ica-flow-config/launchPayload_inputFormValues.json` file.
* The analysis inputs are converted into a "test" profile for Nextflow, stored - among other items - in `nextflow_bench.conf`

<figure><img src="/files/PJ0lkARYoG7v2VoDAyA5" alt=""><figcaption></figcaption></figure>

#### Results

<figure><img src="/files/TGAVfr9SqgvWEkKtKkQt" alt=""><figcaption></figcaption></figure>

```
Enter the number of the entry you want to use: 21
Fetching analysis 9415d7ff-1757-4e74-97d1-86b47b29fb8f ...
Fetching pipeline bb47d612-5906-4d5a-922e-541262c966df ...
Fetching pipeline files... main.nf
Fetching test inputs
New Json inputs detected
Resolving test input ids to /data/mounts/project paths
Fetching input form..
Pipeline "GWAS pipeline_1.
_2_1_20241215_130117" successfully imported.
pipeline name: GWAS pipeline_1_2_1_20241215_130117 
analysis name: Test GWAS pipeline_1_2_1_20241215_130117 
pipeline id : bb47d612-5906-4d5a-922e-541262c966df
analysis id : 9415d7ff-1757-4e74-97d1-86b47b29fb8f
Suggested actions:
pipeline-dev run-in-bench 
I Iterative dev: Make code changes + re-validate with previous command ] 
pipeline-dev deploy-as-flow-pipeline
pipeline-dev run-in-flow
```

## Run Validation Test in Bench

The following command runs this test profile. If a Bench cluster is active, it runs on your Bench cluster, otherwise it runs on the main workspace instance:

```
cd imported-flow-analysis
pipeline-dev run-in-bench
```

{% hint style="info" %}
The pipeline-dev tool is using "nextflow run ..." to run the pipeline. The full nextflow command is printed on stdout and can be copy-pasted+adjusted if you need additional options.
{% endhint %}

<figure><img src="/files/9k1niyiua9gvsz3s351T" alt=""><figcaption></figcaption></figure>

#### Monitoring

When a pipeline is running on **your Bench cluster**, a few commands help to monitor the tasks and cluster. In another terminal, you can use:

* `qstat` to see the tasks being pending or running
* **`tail /data/logs/sge-scaler.log.`**`<latest available workspace reboot time>` to check if the cluster is scaling up or down (it currently takes 3 to 5 minutes to get a new node)

<figure><img src="/files/fV6nCm7ZJNUO7vjdqYg5" alt=""><figcaption></figcaption></figure>

```
/data/demo $ tail /data/logs/sge-scaler.log.*
2025-02-10 18:27:19,657 - SGEScaler - INFO: SGE Marked Overview - {'UNKNOWN': O, 'DEAD': O, 'IDLE': O, 'DISABLED': O, 'DELETED': O, 'UNRESPONSIVE': 0}
2025-02-10 18:27:19,657 - SGEScaler - INFO: Job Status - Active jobs : 0, Pending jobs : 6
2025-02-10 18:27:26,291 - SGEScaler - INFO: Cluster Status - State: Transitioning,
Online Members: 0, Offline Members: 2, Requested Members: 2, Min Members: 0, Max Members: 2
```

#### Data Locations

* The output of the pipeline is in the `outdir` folder
* Nextflow work files are under the `work` folder
* Log files are `.nextflow.log*` and `output.log`

## Modify Pipeline

Nextflow files (located in the `nextflow-src` folder) are easy to modify.\
Depending on your environment (ssh access / docker image with JupyterLab or VNC, with and without Visual Studio code), various source code editors can be used.

```
code nextflow-src # Open in Visual Studio Code
code .            # Open current dir in Visual Studio Code
vi nextflow-src/main.nf
```

After modifying the source code, you can run a validation iteration with the same command as before:

```
pipeline-dev run-in-bench
```

<figure><img src="/files/8aVngUKJaOevZgUEkcFX" alt=""><figcaption></figcaption></figure>

## Identify Docker Image

Modifying the Docker image is the next step.

Nextflow (and ICA) allow the Docker images to be specified at different places:

* in config files such as `nextflow-src/nextflow.config`
* in nextflow code files:

```
/data/demo-flow-dev $ head nextflow-src/main.nf
nextflow.enable.dsl = 2
process top_level_process t
container 'docker.io/ljanin/gwas-pipeline:1.2.1'
```

`grep container` may help locate the correct files:

<figure><img src="/files/TJ5nQAmO1CRrejCHpnT4" alt=""><figcaption></figcaption></figure>

## Docker Image Update: Dockerfile Method

Use case: Update some of the software (mimalloc) by compiling a new version

```
IMAGE_BEFORE=docker.io/ljanin/gwas-pipeline:1.2.1
IMAGE_AFTER=docker.io/ljanin/gwas-pipeline:tmpdemo
 
# Create directory for Dockerfile
mkdir dirForDockerfile
cd dirForDockerfile

# Create Dockerfile
cat <<EOF > Dockerfile
FROM ${IMAGE_BEFORE}
RUN mkdir /mimalloc-compile \
 && cd /mimalloc-compile \
 && git clone -b v2.0.6 https://github.com/microsoft/mimalloc \
 && mkdir -p mimalloc/out/release \
 && cd mimalloc/out/release \
 && cmake ../.. \
 && make \
 && make install \
 && cd / \
 && rm -rf mimalloc-compile
EOF

# Build image
docker build -t ${IMAGE_AFTER} .
```

With the appropriate permissions, you can then "***docker login***" and "***docker push***" the new image.

## Docker Image Update: Interactive Method

```
IMAGE_BEFORE=docker.io/ljanin/gwas-pipeline:1.2.1
IMAGE_AFTER=docker.io/ljanin/gwas-pipeline:1.2.2
docker run -it --rm ${IMAGE_BEFORE} bash
 
# Make some modifications
vi /scripts/plot_manhattan.py
<Fix "manhatten.png" into "manhattAn.png">
<Enter :wq to save and quit vi>
```

```
<Start another terminal (try Ctrl+Shift+T if using wezterm)>
# Identify container id
```

<figure><img src="/files/heowsCTYcf37PSXD9nAn" alt=""><figcaption></figcaption></figure>

```
# Save container changes into new image layer
CONTAINER_ID=c18670335247
docker commit ${CONTAINER_ID} ${IMAGE_AFTER}
```

With the appropriate permissions, you can then "***docker login***" and "***docker push***" the new image.

{% hint style="info" %}
Fun fact: VScode with the "Dev Containers" extension lets you edit the files inside your running container:
{% endhint %}

<figure><img src="/files/FS4VZkhtRZADH505j3km" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
Beware that this extension creates a lot of temp files in /tmp and in $HOME/.vscode-server. Don't include them in your image...
{% endhint %}

Update the nextflow code and/or configs to use the new image

```
sed --in-place "s/${IMAGE_BEFORE}/${IMAGE_AFTER}/" nextflow-src/main.nf
```

Validate your changes in Bench:

```
pipeline-dev run-in-bench
```

## Deploy as Flow Pipeline

```
pipeline-dev deploy-as-flow-pipeline
```

After generating a few ICA-specific files (JSON input specs for Flow launch UI + list of inputs for next step's validation launch), the tool identifies which previous versions of the same pipeline have already been deployed (in ICA Flow, pipeline versioning is done by including the version number in the pipeline name, so that's what is checked here).

It then asks if we want to update the latest version or create a new one.

```
Choice: 2
Creating ICA Flow pipeline dev-nf-core-demo_v4
Sending inputForm.json
Sending onRender.js
Sending main.nf
Sending nextflow.config
```

At the end, the URL of the pipeline is displayed. If you are using a terminal that supports it, Ctrl+click or middle-click can open this URL in your browser.

#### Result

```
/data/demo $ pipeline-dev deploy-as-flow-pipeline

Generating ICA input specs...
Extracting nf-core test inputs...
Deploying project nf-core/demo
- Currently being developed as: dev-nf-core-demo
- Last version updated in ICA:  dev-nf-core-demo_v3
- Next suggested version:       dev-nf-core-demo_v4

How would you like to deploy?
1. Update dev-nf-core-demo (current version)
2. Create dev-nf-core-demo_v4
3. Enter new name
4. Update dev-nf-core-demo_v3 (latest version updated in ICA)
```

```
Sending docs/images/nf-core-demo-subway.svg
Sending docs/images/nf-core-demo_logo_dark.png
Sending docs/images/nf-core-demo_logo_light.png
Sending docs/images/nf-core-demo-subway.png
Sending docs/README. md
Sending docs/output.md

Pipeline successfully deployed
- Id : 26bc5aa5-0218-4e79-8a63-ee92954c6cd9
- URL: https://stage.v2.stratus.illumina.com/ica/projects/1873043/pipelines/26bc5aa5-0218-4e79-8a63-ee92954C6cd9

Suggested actions:
  pipeline-dev run-in-flow
```

## Run Validation Test in Flow

```
pipeline-dev launch-validation-in-flow
```

This launches an analysis in ICA Flow, using the same inputs as the pipeline's "test" profile.

Some of the input files will have been copied to your ICA project to allow the launch to take place. They are stored in the folder `/data/project/bench-pipeline-dev/temp-data`.

#### Result

```
/data/demo $ pipeline-dev launch-validation-in-flow

pipelineld: 26bc5aa5-0218-4e79-8a63-ee92954c6cd9
Getting Analysis Storage Id
Launching as ICA Flow Analysis...

ICA Analysis created:
- Name: Test dev-nf-core-demo_v4
- Id:   cadcee73-d975-435d-b321-5d60e9aec1ec
- Url:   https://stage.v2.stratus.illumina.com/ica/projects/1873043/analyses/cadcee73-d975-435d-b321-5d60e9aec1ec
```

<figure><img src="/files/MgD6DtdtLPWJYBNjkkov" alt=""><figcaption></figcaption></figure>


# Containers in Bench

Bench has the ability to handle containers inside a running workspace.\
This allows you to install and package software more easily as a container image and provides capabilities to pull and run containers inside a workspace.

Bench offers a **container runtime as a service** in your running workspace. This allows you to do standardized container operations such as pulling in images from public and private registries, build containers at runtime from a Dockerfile, run containers and eventually publish your container to a registry of choice to be used in different Platform Core products such as Platform Core Flow.

## Setup

The Container Service is **accessible from your Bench workspace** environment by default.

The container service uses the **workspace disk** to store any container images you pulled in or created.

To interact with the Container Service, a **container remote client CLI** is exposed automatically in the `/data/.local/bin` folder. The Bench workspace environment is preconfigured to automatically detect where the Container Service is made available using environment variables. These environment variables are automatically injected into your environment and are not determined by the Bench Workspace Image.

## Container Management

Use either *docker* or *podman* cli to interact with the Container Service. Both are interchangeable and support all the standardized operations commonly known.

### Pulling a Container Image

To run a container, the first step is to either build a container from a source container or pull in a container from a registry.

#### Public Registry

A public image registry does not require any form of authentication to pull the container layers.

The following command line example shows how to pull in a commonly known image.

{% hint style="info" %}
The Container Service uses Dockerhub by default to pull images from if no registry hostname is defined in the container image URI.
{% endhint %}

```
# Pull Container image from Dockerhub 
/data $ docker pull alpine:latest  
```

#### Private Registry

To pull images from a private registry, the Container Service needs to authenticate to the Private Registry.

The following command line example shows how to instruct the Container Service to login into the Private `registry.hub.docker.com` registry

```
# Pull a Container Image from Dockerhub 
/data $ docker login -u <username> registry.hub.docker.com 
Password:  
Login Succeeded! 
/data $ docker pull registry.hub.docker.com/<privateContainerUri>:<tag> 
```

{% hint style="info" %}
Depending on your authorisations in the private registry you will be able to pull and push images. These authorisations are managed outside of the scope of Platform Core.
{% endhint %}

### Pushing a Container Image

Depending on the Registry setup you can publish Container Images with or without authentication.\
\
If Authentication is required, follow the login procedure described in [Private Registry](#private-registry)

The following command line example shows how to publish a locally available Container Image to a private registry in Dockerhub.

```
# Push a Container Image to a Private registry in Dockerhub 
/data $ docker pull alpine:latest 
/data $ docker tag alpine:latest registry.hub.docker.com/<privateContainerUri>:<tag> 
/data $ docker push registry.hub.docker.com/<privateContainerUri>:<tag> 
```

### Saving a Container Image as an Archive

The following example shows how to save a locally available Container Image as a compressed tar archive.

```
# Save a Container Image as a compressed archive 
/data $ docker pull alpine:latest 
/data $ docker save alpine:latest | bzip2 > /data/alpine_latest.tar.bz2 
```

This lets you upload the [container image](/home/h-dockerrepository) into the Private Platform Core Docker Registry.

### Listing Locally Available Container Images

The following example shows how to list all locally available Container Images

```
# List all local available images 
/data $ docker images 
REPOSITORY                TAG         IMAGE ID      CREATED      SIZE 
docker.io/library/alpine  latest      aded1e1a5b37  3 weeks ago  8.13 MB 
```

### Deleting a Container Image

Container Images require storage capacity on the Bench Workspace disk. The capacity is shown when listing the locally available container images. The container Images are persisted on disk and remain available whenever a workspace stops and restarts.

The following example shows how to clean up a locally available Container Image:

```
# Remove a locally available image 
/data $ docker rmi alpine:latest 
```

{% hint style="info" %}
When a Container Image has multiple tags, all the tags need to be removed individually to free up disk capacity.
{% endhint %}

### Running a Container

A Container Image can be instantiated in a Container running inside a Bench Workspace.

By default the workspace disk (`/data`) will be made available inside the running Container. This lets you to access data from the workspace environment.

When running a Container, the default user defined in the Container Image manifest will be used and mapped to the uid and the gid of the user in the running Bench Workspace (uid:1000, gid: 100). This will ensure files created inside the running container on the workspace disk will have the same file ownership permissions.

#### Run a Container as a normal user

The following command line example shows how to run an instance a locally available Container Image as a normal user:

```
# Run a Container as a normal user 
/data $ docker run -it --rm alpine:latest 
~ $ id 
uid=1000(ica) gid=100(users) groups=100(users)  
```

#### Run a Container as root user

Running a Container as root user maps the uid and gid inside the running Container to the running non-root user in the Bench Workspace. This lets you act as user with uid 0 and gid 0 inside the context of the container.

By enabling this functionality, you can install system level packages inside the context of the Container. This can be leveraged to run tools that require additional system level packages at runtime.

The following command line example shows how to run an instance of a locally available Container as root user and install system level packages.

```
# Run a Container as root user 
/data $ docker run -it --rm --userns keep-id:uid=0,gid=0 --user 0:0 alpine:latest 
/ # id 
uid=0(root) gid=0(root) groups=0(root) 
/ # apk add rsync 
... 
/ # rsync  
rsync  version 3.4.0  protocol version 32 
... 
```

When no specific mapping is defined using the `--userns` flag, the user in the running Container user will be mapped to an undefined uid and gid based on an offset of id 100000. Files created in your workspace disk from the running Container will also use this uid and gid to define the ownership of the file.

```
# Run a Container as a non-mapped root user 
/data $ docker run -it --rm --user 0:0 alpine:latest 
/ # id 
uid=0(root) gid=0(root) groups=100(users),0(root) 
/ # touch /data/myfile 
/ #  
# Exited the running Container back to the shell in the running Bench Workspace 
/data $ ls -al /data/myfile  
-rw-r--r-- 1 100000 100000 0 Mar 13 08:27 /data/myfile 
```

**Building a Container**

To build a Container Image, you need to describe the instructions in a Dockerfile.

This next example builds a local Container Image and tags it as myimage:1.0 The Dockerfile used in this example is

```
FROM alpine:latest 
RUN apk add rsync 
COPY myfile /root/myfile 
```

The following command line example will build the actual Container Image.

```
# Build a Container image locally 
/data $ mkdir /tmp/buildContext 
/data $ touch /tmp/buildContext/myFile 
/data $ docker build -f /tmp/Dockerfile -t myimage:1.0 /tmp/buildContext 
... 
/data $ docker images 
REPOSITORY                TAG         IMAGE ID      CREATED             SIZE 
docker.io/library/alpine  latest      aded1e1a5b37  3 weeks ago         8.13 MB 
localhost/myimage         1.0         06ef92e7544f  About a minute ago  12.1 MB 
```

{% hint style="info" %}
When defining the build context location, keep in mind that using the HOME folder (/data) will index all files available in /data, which can be a lot and will slow down the process of building. Hence the reason to use a minimal build context whenever possible.
{% endhint %}


# FUSE Driver

Bench Workspaces use a FUSE driver to mount project data directly into a workspace file system. There are both read and write capabilities with some limitations on write capabilities that are enforced by the underlying AWS S3 storage.

As a user, you are allowed to do the following actions from Bench (when having the correct user permissions compared to the workspace permissions) or through the CLI:

* Copy project data
* Delete project data
* Mount project data (CLI only)
* Unmount project data (CLI only)

When you have a running workspace, you will find a file system in Bench under the project folder along with the basic and advanced tutorials. When opening that folder, you will see all the data that resides in your project.

{% hint style="danger" %}
**This is a fully mounted version of the project data. Changes in the workspace to project data cannot be undone.**
{% endhint %}

## Copy project data

The FUSE driver allows the user to easily copy data from /data/project to the local workspace and vice versa.\
There is a file size limit of 500 GB per file for the FUSE driver.

## Delete project data

The FUSE driver also allows you to delete data from your project. This is different from the use of Bench before where you took a local copy and still kept the original file in your project.

{% hint style="danger" %}
**Deleting project data through Bench workspace through the FUSE driver will permanently delete the data in the Project. This action cannot be undone.**
{% endhint %}

## CLI

Using the FUSE driver through the CLI is not supported for Windows users. Linux users will be able to use the CLI without any further actions, Mac users will need to install the kernel extension from [macFuse](https://osxfuse.github.io/).

{% hint style="info" %}
MacOS uses hidden metadata files beginning with .\_ ,which are copied over and exposed during CLI copy to your project data. These can be safely deleted from your project.
{% endhint %}

Mount and unmount of data can be done in the CLI with the commands `icav2 projectdata mount mnt--project-id <your_project_uuid>` and `icav2 projectdata unmount`. **In Bench this happens automatically** and the CLI commands are not needed.

{% hint style="danger" %}
Do NOT use the CP -f command to copy or move data to a mounted location. This will result in data loss as data on the destination location will be deleted.
{% endhint %}

## Restrictions

{% hint style="danger" %}
Once a file is written, it cannot be changed! You **will not be able to update it** in the project location because of the restrictions mentioned above.
{% endhint %}

Trying to update files or saving you notebook in the project folder will typically result in `File Save Error for fusedrivererror.ipynb Invalid response: 500 Internal Server Error.`

Some examples of other actions or commands that will not work because of the above mentioned limitations:

* Save a jupyter notebook or R script on the /project location
* Add/remove a file from an existing zip file
* Redirect with append to an existing file e.g. echo "This will not work" >> myTextFile.txt
* Rename a file due to the existing association between Platform Core and AWS
* Move files or folders.
* Using vi or another editor

A file can be written only sequentially. This is a restriction that comes from the library the FUSE driver uses to store data in AWS. That library supports only sequential writing, random writes are currently not supported. The FUSE driver will detect random writes and the write will fail with an IO error return code. Zip will not work since zip writes a table of contents at the end of the file. Please use gzip.

Listing data (ls -l) reads data from the platform. The actual data comes from AWS and there can be a short delay between the writing of the data and the listing being up to date. As a result, a file that is written may appear temporarily as a zero length file, a file that is deleted may appear in the file list. This is a tradeoff, the FUSE driver caches some information for a limited time and during that time the information may seem wrong. Note that besides the FUSE driver, the library used by the FUSE driver to implement the raw FUSE protocol and the OS kernel itself may also do caching.

### Jupyter notebooks

To use a specific file in a jupyter notebook, you will need to use **'/data/project/filename'**.

## Old Bench workspaces

This functionality won't work for old workspaces unless you enable the permissions for that old workspace.


# Run DRAGEN in Bench - Interactive

## Introduction

DRAGEN can run in [Bench](/project/p-bench) workspaces

* **On FPGA-instances**, DRAGEN can run in **FPGA mode** (hardware-accelerated) or **software mode**. This can be useful when comparing performance gains by hardware acceleration or to distribute concurrent processes between the FPGA and cpu.
* On **non-FPGA** instances DRAGEN can only run in **software mode.**

{% hint style="info" %}
To run DRAGEN in software mode, you need to use the DRAGEN `--sw-mode` parameter.
{% endhint %}

The DRAGEN command line parameters to specify the location of the licence file are different.

**FPGA** mode uses

{% code overflow="wrap" %}

```
LICENSE_PARAMS="--lic-instance-id-location /opt/dragen-licence/instance-identity.protected --lic-credentials /opt/dragen-licence/instance-identity.protected/dragen-creds.lic"
```

{% endcode %}

**Software** mode uses

{% code overflow="wrap" %}

```
LICENSE_PARAMS="--sw-mode --lic-credentials /opt/dragen-licence/instance-identity.protected/dragen-creds-sw-mode.lic"
```

{% endcode %}

## DRAGEN Bench Images

DRAGEN software is provided in specific Bench images with names starting with `Dragen`. For example (available versions may vary):

* `Dragen 4.4.1 - Minimal` provides DRAGEN 4.4.1 and SSH access
* `Dragen 4.4.6` provides DRAGEN 4.4.6, SSH and JupyterLab.

### Prerequisites

#### Memory

The instance type is selected during workspace creation (**Projects > your\_project > Bench > Workspaces**). The amount of RAM available on the instance is critical. **256GiB RAM** is a safe choice to run DRAGEN in production. **All FPGA2 instances offer 256GiB or more of RAM**.

When running in **Software mode**, use[ **himem-large**](/reference/r-pricing#compute) (348GiB RAM) or [**hicpu-large**](/reference/r-pricing#compute) (144 GiB RAM) to ensure enough RAM is available for your runs.

{% hint style="info" %}
During pipeline development, when typically using small amounts of data, you can try to scale down in instance types to save costs. You can start at hicpu-large and progressively use smaller instances, though you will need at least **standard-xlarge**.\
**If DRAGEN runs out of available memory, the system is rebooted**, losing your currently running commands and interface.

**DRAGEN version 4.4.6** and later verify if the system has at least **128GB** of memory available. If not enough memory is available, you will encounter an error stating that the *Available memory is less than the minimum system memory required 128GB.*\
This can be overridden with the command line parameter `dragen --min-memory 0`
{% endhint %}

## FPGA-mode

Using an fpga2-medium [instance type](/project/p-flow/f-pipelines#compute-types).

#### Example

{% code overflow="wrap" %}

```sh
mkdir /data/demo 
cd /data/demo 

# download ref 
wget --progress=dot:giga https://s3.amazonaws.com/stratus-documentation-us-east-1-public/dragen/reference/Homo_sapiens/hg38.fa -O hg38.fa 
# => 0.5min 

# Build ht-ref 
mkdir ref 
dragen --build-hash-table true --ht-reference hg38.fa --output-directory ref 
# => 6.5min 

# run DRAGEN mapper 
FASTQ=/opt/edico/self_test/reads/midsize_chrM.fastq.gz

# Next line is needed to resolve "run the requested pipeline with a pangenome reference, but a linear reference was provided" in DRAGEN (4.4.1 and others). Comment out when encountering unrecognised option '--validate-pangenome-reference=false'.
DRAGEN_VERSION_SPECIFIC_PARAMS="--validate-pangenome-reference=false" 

# License Parameters
LICENSE_PARAMS="--lic-instance-id-location /opt/dragen-licence/instance-identity.protected --lic-credentials /opt/dragen-licence/instance-identity.protected/dragen-creds.lic"

mkdir out
dragen -r ref --output-directory out --output-file-prefix out -1 $FASTQ --enable-variant-caller false --RGID x --RGSM y ${LICENSE_PARAMS} ${DRAGEN_VERSION_SPECIFIC_PARAMS} 
# => 1.5min (10 sec if fpga already programmed)
```

{% endcode %}

## Software-mode

Using a standard-xlarge [instance type](/project/p-flow/f-pipelines#compute-types).

Software mode is activated with the DRAGEN `--sw-mode` parameter.

#### Example

{% code overflow="wrap" %}

```sh
mkdir /data/demo 
cd /data/demo 

# download ref 
wget --progress=dot:giga https://s3.amazonaws.com/stratus-documentation-us-east-1-public/dragen/reference/Homo_sapiens/hg38.fa -O hg38.fa 
# => 0.5min 

# Build ht-ref 
mkdir ref 
dragen --build-hash-table true --ht-reference hg38.fa --output-directory ref 
# => 6.5min 

# run DRAGEN mapper 
FASTQ=/opt/edico/self_test/reads/midsize_chrM.fastq.gz

# Next line is needed to resolve "run the requested pipeline with a pangenome reference, but a linear reference was provided" in DRAGEN (4.4.1 and others). Comment out when encountering ERROR: unrecognised option '--validate-pangenome-reference=false'.
DRAGEN_VERSION_SPECIFIC_PARAMS="--validate-pangenome-reference=false"

# When using DRAGEN 4.4.6 and later, the line above should be extended with --min-memory 0 to skip the memory check.
DRAGEN_VERSION_SPECIFIC_PARAMS="--validate-pangenome-reference=false --min-memory 0" 

# License Parameters
LICENSE_PARAMS="--sw-mode --lic-credentials /opt/dragen-licence/instance-identity.protected/dragen-creds-sw-mode.lic" 

mkdir out 
dragen -r ref --output-directory out --output-file-prefix out -1 $FASTQ --enable-variant-caller false --RGID x --RGSM y ${LICENSE_PARAMS} ${DRAGEN_VERSION_SPECIFIC_PARAMS} 
# => 2min
```

{% endcode %}


# Cohorts

## Introduction to Cohorts

Cohorts is a cohort analysis tool integrated with Illumina BioInsight Platform Core. Cohorts combines subject- and sample-level metadata, such as phenotypes, diseases, demographics, and biometrics, with molecular data stored in Platform Core to perform tertiary analyses on selected subsets of individuals.

## Overview Video

This video is an overview of Illumina BioInsight Platform Core. It walks through a Multi-Omics Cancer workflow that can be found here: [Oncology Walkthrough](https://help.ica.illumina.com/project/p-cohorts/cohorts-walkthrough-cancer)

{% embed url="<https://www.youtube.com/watch?v=wr19Y8BVaQ4&ab_channel=Illumina>" %}
Illumina BioInsight Platform Core: Cohorts Multi-Omic Cancer
{% endembed %}

### Features At-a-glance

* Intuitive UI for selecting subjects and samples to analyze and compare: deep phenotypical and clinical metadata, molecular features including germline, somatic, gene expression.
* Comprehensive, harmonized data model exposed to Platform Core Base and Platform Core Bench users for custom analyses.
* Run analyses in Platform Core Base and Platform Core Bench and upload final results back into Cohorts for visualization.
* Out-of-the-box statistical analyses including genetic burden tests, GWAS/PheWAS.
* Rich public data sets covering key disease areas to enrich private data analysis.
* Easy-to-use visualizations for gene prioritization and genetic variation inspection.

## Functionality

* [Create a Cohort](/project/p-cohorts/cohorts-create)
* [Import New Samples](/project/p-cohorts/cohorts-import)
* [Prepare Metadata Sheets](/project/p-cohorts/cohorts-metadata)
* [Precomputed GWAS and PheWAS](/project/p-cohorts/cohorts-gwas-phewas)
* [Cohort Analysis](/project/p-cohorts/cohorts-analysis)
* [Compare Cohorts](/project/p-cohorts/cohorts-comparison)
* [Cohorts Data in Platform Core Base](/project/p-cohorts/cohorts-base)

## Walk-throughs

* [Oncology](/project/p-cohorts/cohorts-walkthrough-cancer)
* [Rare Genetic Disorders](/project/p-cohorts/cohorts-walkthrough-raredisease)

## Public Data Sets

Cohorts contains a variety of freely available data sets covering different disease areas and sequencing technologies. For a list of currently available data, [see here](/project/p-cohorts/cohorts-publicdata).


# Create a Cohort

Platform Core Cohorts lets you create a research cohort of subjects and associated samples based on the following criteria:

* Project:
  * Include subjects that are part of any Platform Core project that you own or that is shared with you.
  * Sample:
    * Sample type such as FFPE.
    * Tissue type.
    * Sequencing technology: Whole genome DNA-sequencing, RNAseq, single-cell RNAseq, etc.
* Subject:
  * Subject inclusion by Identifier:
    * Input a list of Subject Identifiers (up to 100 entries) when defining a cohort.
    * The Subject Identifier filter is combined using AND logic with any other applied filters.
    * Within the list of subject identifiers, OR logic is applied (i.e., a subject matches if it is in the provided list).
  * Demographics such as age, sex, ancestry.
  * Biometrics such as body height, body mass index.
  * Family and patient medical history.
* Sample:
  * Sample type such as FFPE.
  * Tissue type.
  * Sequencing technology: Whole genome DNA-sequencing, RNAseq, single-cell RNAseq, etc.
* Disease:
  * Phenotypes and diseases from standardized ontologies.
* Drug:
  * Drugs from standardized ontologies along with specific typing, stop reasons, drug administration routes, and time points.
* Molecular attributes:
  * Samples with a somatic mutation in one or multiple, specified genes.
  * Samples with a germline variant of a specific type in one or multiple, specified genes.
  * Samples over- or under-expressed in one or multiple, specified genes.
  * Samples with a copy number gain or loss involving one or multiple, specified genes.

### Disease search

Platform Core Cohorts currently uses six standard medical ontologies to 1) annotate each subject during ingestion and then to 2) search for subjects: HPO for phenotypes, MeSH, SNOMED-CT, ICD9-CM, ICD10-CM, and OMIM for diseases. By default, any type-ahead search will find matches from all six, and you can limit the search to only the ones you prefer. When searching for subjects using names or codes from one of these ontologies, Platform Core Cohorts will automatically match your query against all the other ontologies, therefore returning subjects that have been ingested using a corresponding entry from another ontology.

In the **Disease** tab, you can search for subjects diagnosed with one or multiple diseases, as well as phenotypes, in two ways:

* **Start typing the English name** of a disease/phenotype and pick from the suggested matches. Continue typing if your disease/phenotype of interest is not listed initially.
  * Use the mouse to select the term or navigate to the term in the dropdown using the arrow buttons.
  * If applicable, the concept hierarchy is shown, with ancestors and immediate children visible.
  * For diagnostic hierarchies, concept children count and descendant count for each disease name is displayed.
    * **Descendant Count**: Displays next to each disease name in the tree hierarchy (e.g., "Disease (10)").
    * **Leaf Nodes**: No children count shown for leaf nodes.
    * **Missing Counts**: Children count is hidden if unavailable.
    * **Show Term Count**: A new checkbox below "Age of Onset" that is always checked. Unchecking it hides the descendant count.
  * Select a checkbox to include the diagnostic term along with all of its children and decedents.
  * Expand the categories and select or deselect specific disease concepts.
* **Paste one or multiple diagnostic codes** separated by a pipe (‘|’).

### Drug Search

In the **Drug** tab, you can search for subjects who have a specific medication record:

* Start typing the concept name for the drug and pick from suggested matches. Continue typing if the drug is not listed initially.
* Paste one or multiple drug concept codes. Platform Core Cohorts currently uses RXNorm as a standard ontology during ingestion. If multiple concepts are in your instance of Platform Core Cohorts, they will be listed under **Concept Ontology**.
* 'Drug Type' is a static list of qualifiers that denote the specific administration of the drug. For example, where the drug was dispensed.
* 'Stop Reason' is a static list of attributes describing a reason why a drug was stopped if available in the data ingestion.
* 'Drug Route' is a static list of attributes that describe the physical route of administration of the drug. For example, Intravenous Route (IV).

### Measurement Search

In the ‘**Measurements**’ tab, you can search for vital signs and laboratory test data leveraging LOINC concept codes. ·

* Start typing the English name of the LOINC term, for example, ‘Body height’. A dropdown will appear with matching terms. Use the mouse or down arrows to select the term.
* Upon selecting a term, the term will be available for use in a query.
* Terms can be added to your query criteria.
* For each term, you can set a value **Greater than or equal**, **Equals**, **Less than or equal**, **In range**, or **Any value**.
* **Any value** will find any record where there is an entry for the measurement independent of an available value.
* Click **Apply** to add your criteria to the query.
* Click **Update Now** to update the running count of the Cohort.Include/Exclude

### Include/Exclude

* As attributes are added to the 'Selected Condition' on the right-navigation panel, you can choose to include or exclude the criteria selected.
  * Select a criterion from 'Subject', 'Disease', and/or 'Molecular' attributes by filling in the appropriate checkbox on the respective attribute selection pages.
  * When selected, the attribute will appear in the right-navigation panel.
  * You can use the 'Include' / 'Exclude' dropdown next to the selected attribute to decide if you want to include or exclude subjects and samples matching the attribute.

{% hint style="info" %}
The semantics of 'Include' work in such a way that a subject needs to match only one or multiple of the 'included' attributes in any given category to be included in the cohort. (Category refers to disease, sex, body height, etc.) For example, if you specify multiple diseases as inclusion criteria, subjects will only need to be diagnosed with one of them. Using 'Exclude', you can exclude any subject who matches one or multiple exclusion criteria; subjects do not have to match all exclusion criteria in the same category to be excluded from the cohort.

This feature is not available on the 'Project' level selections as there is no overlap between subjects in datasets.

Using exclusion criteria does not account for NULL values. For example, if the Super-population 'Europeans' is excluded, subjects will be in your cohort even if they do not contain this data point.
{% endhint %}

Once you selected **Create Cohort**, the above data are organized in tabs such as Project, Subject, Disease, and Molecular. Each tab then contains the aforementioned sections, among others, to help you identify cases and/or controls for further analysis. Navigate through these tabs, or search for an attribute by name to directly jump to that tab and section, and select attributes and values that are relevant to describe your subjects and samples of interest. Assign a new name to the cohort you created, and click **Apply** to save the cohort.

## Duplicate a Cohort Definition

* After creating a Cohort, select the **Duplicate** icon.
* A copy of the Cohort definition will be created and tagged with "`_copy`**"**.

## Delete a Cohort Definition

* Deleting a Cohort Definition can be accomplished by clicking the **Delete Cohort** icon.
* *This action cannot be undone.*

## Sharing a Cohort within a Platform Core Project

After creating a Cohort, users can set a Cohort bookmark as Shared. By sharing a Cohort, the Cohort will be available to be applied across the project by other users with access to the Project. Cohorts created in a Project are only accessible at scope of the user. Other users in the project cannot see the cohort created unless they use this sharing functionality.

### Share Cohort Definition

* Create a Cohort using the directions above.
* To make the Cohort available to other users in your Project, click the **Share** icon.
* The **Share** icon will be filled in black and the Shared Status will be turned from **Private** to **Shared**.
* Other users with access to Cohorts in the Project can now apply the Cohort bookmark to their data in the project.

### Unshare a Cohort Definition

* To unshare the Cohort, click the **Share** icon.
* The icon will turn from black to white, and other users within the project will no longer have access to this cohort definition.

### Archive a Cohort Definition

* A Shared Cohort can be Archived.
* Select a Shared Cohort with a black **Shared Cohort** icon.
* Click the **Archive Cohort** icon.
* You will be asked to confirm this selection.
* Upon archiving the Cohort definition, the Cohort will no longer be seen by other users in the Project.
* The archived Cohort definition can be unarchived by clicking the **Unarchive Cohort** icon.
* When the Cohort definition is unarchived, it will be visible to all users in the Project.

### Sharing a Cohort as Bundle

You can link cohorts data sets to a bundle as follows:

* Create or edit a bundle at **Bundles** from the main navigation.
* Navigate to **Bundles > your\_bundle > Cohorts > Data Sets**.
* Select **Link Data Set to Bundle**.
* Select the data set which you want to link and **+Select**.
* After a brief time, the cohorts data set will be linked to your bundle and ICA\_BASE\_100 will be logged.

{% hint style="info" %}
If you can not find the cohorts data sets which you want to link, verify if

* Your data set is part of a project (**Projects > your\_project > Cohorts > Data Sets**)
* This project is set to Data Sharing (**Projects > your\_project > Project Settings > Details)**
  {% endhint %}

### Stop sharing a Cohort as Bundle

You can unlink cohorts data sets from bundles as follows:

* Edit the desired bundle at **Bundles** from the main navigation.
* Navigate to **Bundles > your\_bundle > Cohorts > Data Sets**.
* Select the cohorts data set which you want to unlink.
* Select **Unlink Data Set from Bundle**.
* After a brief time, the cohorts data set will be unlinked from your bundle and ICA\_BASE\_101 will be logged.


# Import New Samples

## Import New Samples

Platform Core Cohorts can pull any molecular data available in a Platform Core project, as well as additional sample- and subject-level metadata information such as demographics, biometrics, sequencing technology, phenotypes, and diseases.

To import a new data set, select **Import Jobs** from the left navigation tab underneath **Cohorts**, and click the **Import Files** button. The **Import Files** button is also available under the **Data Sets** left navigation item.

{% hint style="info" %}
The **Data Set** menu item is used to view imported data sets and information. The **Import Jobs** menu item is used to check the status of data set imports.
{% endhint %}

Confirm that the project shown is the Platform Core project that contains the molecular data you would like to add to Platform Core Cohorts.

1. Choose a data type among
   * Germline variants
   * Somatic mutations
   * RNAseq
   * GWAS
2. Choose a new study name by selecting the radio button **Create new study** and entering a **Study Name**.
3. To add new data to an existing Study, select the radio button **Select from list of studies** and select an existing **Study Name** from the dropdown.
4. To add data to existing records or add new records, select **Job Type**, **Append**.
5. **Append** does not wipe out any data ingested previously and can be used to ingest the molecular data in an incremental manner.
6. To replace data, select **Job Type**, **Replace**. If you are ingesting data again, use the Replace job type.
7. Enter an optional **Study description**.
8. Select the metadata model (default: Cohorts; alternatively, select OMOP version 5.4 if your data is formatted that way.)
9. Select the **genome build** your molecular data is aligned to (default: GRCh38/hg38)
10. For RNAseq, specify whether you want to run differential expression (see below) or only upload raw TPM.
11. Click **Next**.
12. Navigate to VCFs located in the Project Data.
13. Select each single-sample VCF or multi-sample VCF to ingest. For GWAS, select CSV files produced by Regenie.
    * As an alernative to selecting individual files, you can also opt to select a folder instead. Toggle the radio button on Step 2 from "Select files" to "Select folder".
    * This option is currently only available for germline variant ingestion: any combination of small variants, structural variation, and/or copy number variants.
    * Platform Core Cohorts will scan the selected folder and all sub-folders for any VCF files or JSON files and try to match them against the Sample ID column in the metadata TSV file in Step 3.
    * Files not matching sample IDs will be ignored; allowed file extensions for VCF files after the sample ID are: \*.vcf.gz, \*.hard-filtered.vcf.gz, \*.cnv.vcf.gz, and \*.sv.vcf.gz .
    * Files not matching sample IDs will be ignored; allowed file extensions for JSON files after the sample ID are: *.json,*.json.gz, \*.json.bgz, \*.json.gzip.
14. Click **Next**.
15. Navigate to the metadata (phenotype) data *tsv* in the project Data.
16. Select the TSV file or files for ingestion.
17. Click **Finish**.

{% hint style="info" %}
Search Spinner behavior in input jobs table

* Search a term and press \*\* Enter.
* The search spinner will appear while the results are loading.
* Once the results are displayed in the table, the spinner will disappear immediately
  {% endhint %}

{% hint style="info" %}
All VCF types, specifically from DRAGEN, can be ingested using the Germline variants selection. Cohorts will distinguish the variant types that it is ingesting. If Cohorts cannot determine the variant file type, it will default to ingest small variants.

Alternatively to VCFs, you can select Nirvana JSON files for DNA variants: small variants, structural variants, and copy number variation.

The maximum amount of files that can be part of a single manual ingestion batch is capped at 1000

You can also choose a single folder and Platform Core Cohorts will identify all ingestible files within that folder and its sub-folders. In this scenario, Cohorts will select molecular data files matching the samples listed in the metadata sheet, which is the next step in the import process.

You have the option to ingest either VCF files or Nirvana JSON files for any given batch, regardless of the chosen ingestion method.

The sample identifiers used in the VCF columns need to match the sample identifiers used in subject/sample metadata files; accordingly, if you are starting from JSON files containing variant- and gene-level annotations provided by ILMN Nirvana, the samples listed in the header need to match the metadata files.
{% endhint %}

#### Variant file formats

Platform Core Cohorts supports VCF files formatted according to VCF v4.2 and v4.3 specifications. VCF files require at least one of the following header rows to identify the genome build:

* \##reference=file://... --- needs to contain a reference to hg38/GRCh38 in the file path or name (numerical value is sufficient)
* \##contig=\<ID=chr1,length=248956422> --- for hg38/GRCh38
* \##DRAGENCommandLine= ... --ht-reference

Platform Core Cohorts accepts VCFs aligned to hg38/GRCh38 and hg19/GRCh37. If your data uses hg19/GRCh37 coordinates, Cohorts will convert these to hg38/GRCh38 during the ingestion process \[see Reference 1]. Harmonizing data to one genome build facilitates searches across different private, shared, and public projects when building and analyzing a cohort. If your data contains a mixture of samples mapped to hg38 and hg19, please ingest these in separate batches, as each import job into Cohorts is limited to one genome build.

As an alternative to VCFs, Platform Core Cohorts accepts the JSON output of [Illumina Nirvana](https://illumina.github.io/NirvanaDocumentation/) for hg38/GRCh38-aligned data for small germline variants and somatic mutations, copy number variations and other structural variants.

#### RNAseq file format

Platform Core Cohorts can process gene- and transcript-level quantification files produced by the Illumina DRAGEN RNA pipeline. The file naming convention needs to match `.quant.genes.sf` for genes and `.quant.sf` for transcript-level TPM.

Please also see the online documentation for the [Illumina DRAGEN RNA Pipeline](https://support-docs.illumina.com/SW/dragen_v42/Content/SW/DRAGEN/GeneExpressionQuantification.htm) for more information on output file formats.

#### GWAS file format

Platform Core Cohorts currently supports upload of SNV-level GWAS results produced by [Regenie](https://rgcgithub.github.io/regenie/) and saved as CSV files.

### Metadata and File Types

<table data-header-hidden><thead><tr><th width="204.6336669921875"></th><th></th></tr></thead><tbody><tr><td><strong>Field</strong></td><td><strong>Description</strong></td></tr><tr><td>Project name</td><td>The Platform Core project for your cohort analysis (cannot be changed.)</td></tr><tr><td>Study name</td><td>Create or select a study. Each study represents a subset of data within the project.</td></tr><tr><td>Description</td><td>Short description of the data set (optional).</td></tr><tr><td>Job type</td><td><strong>Append</strong>: Appends values to any existing values. If a field supports only a single value, the value is replaced.</td></tr><tr><td></td><td><strong>Replace</strong>: Overwrites existing values with the values in the uploaded file.</td></tr><tr><td>Subject metadata files</td><td>Subject metadata file(s) in tab-delimited format.<br>For <strong>Append</strong> and <strong>Replace</strong> job types, the following fields are required and cannot be changed:<br>- Sample identifier<br>- Sample display name<br>- Subject identifier<br>- Subject display name<br>- Sex</td></tr></tbody></table>

{% hint style="info" %}
If annotating large sets of samples with molecular data, expect the annotation process to take over 20 minutes per whole genome batch of samples. You will receive two e-mail notifications: once your ingestion starts and once completed successfully or failed.
{% endhint %}

As an alternative to Platform Core Cohorts metadata file format, you can provide files formatted according to the [OMOP common data model 5.4](http://ohdsi.github.io/CommonDataModel/cdm54.html). Cohorts currently ingests data for these OMOP 5.4 tables, formatted as tab-delimited files:

* PERSON (mandatory),
* CONCEPT (mandatory if any of the following is provided),
* CONDITION\_OCCURRENCE (optional),
* DRUG\_EXPOSURE (optional), and
* PROCEDURE\_OCCURRENCE (optional.)

Additional files such as measurement and observation will be supported in a subsequent release of Cohorts.

{% hint style="info" %}
Cohorts requires that all such files do not deviate from the OMOP CDM 5.4 standard. Depending on your implementation, you may have to adjust file formatting to be OMOP CDM 5.4-compatible.
{% endhint %}

## References

\[1] VcfMapper: <https://stratus-documentation-us-east-1-public.s3.amazonaws.com/downloads/cohorts/main_vcfmapper.py>

\[2] crossMap: <https://crossmap.sourceforge.net/>

\[3] liftOver: <https://genome.ucsc.edu/cgi-bin/hgLiftOver>

\[4] Chain files: <ftp://ftp.ensembl.org/pub/assembly_mapping/homo_sapiens/>


# Prepare Metadata Sheets

In Platform Core Cohorts, metadata describe any subjects and samples imported into the system in terms of attributes, including:

* subject:
  * demographics such as age, sex, ancestry;
  * phenotypes and diseases;
  * biometrics such as body height, body mass index, etc.;
  * pathological classification, tumor stages, etc.;
  * family and patient medical history;
* sample:
  * sample type such as FFPE,
  * tissue type,
  * sequencing technology: whole genome DNA-sequencing, RNAseq, single-cell RNAseq, among others.

You can use these attributes while [creating](/project/p-cohorts/cohorts-create) a cohort to define the cases and/or controls that you want to include.

During [import](/project/p-cohorts/cohorts-import), you will be asked to upload a metadata sheet as a tab-delimited (TSV) file. An example sheet is available for download on the **Import files** page in the Platform Core Cohorts UI.

A metadata sheet will need to contain at least these four columns per row:

* **Subject ID** - identifier referring to individuals; use the column header "SubjectID".
* **Sample ID** - identifier for a sample. Sample IDs need to match the corresponding column header in VCF/GVCFs; each subject can have multiple samples, these need to be specified in individual rows for the same **SubjectID**; use the column header "SampleID".
* **Biological sex** - can be "Female (XX)", "Female"; "Male (XY)", "Male"; "X (Turner's)"; "XXY (Klinefelter)"; "XYY"; "XXXY" or "Not provided". Use the column header "DM\_Sex" (demographics).
* **Sequencing technology** - can be "Whole genome sequencing", "Whole exome sequencing", "Targeted sequencing panels", or "RNA-seq"; use the column header "TC" (technology).

A description of all attributes and data types currently supported by Platform Core Cohorts can be found here: [ICA\_Cohorts\_Supported\_Attributes.xlsx](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/downloads/cohorts/ICA_Cohorts_Supported_Attributes.xlsx)

You can download an example of a metadata sheet, which contains some samples from The Cancer Genome Atlas ([TCGA](https://www.cancer.gov/ccg/research/genome-sequencing/tcga)) and their publicly available clincal attributes, here: [ICA\_Cohorts\_Example\_Metadata.tsv](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/downloads/cohorts/ICA_Cohorts_Example_Metadata.tsv)

A list of concepts and diagnoses that cover all public data subjects to easily navigate the new concept code browser for diagnosis can be found here: [PublicData\_AllConditionsSummarized.xlsx](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/downloads/cohorts/PublicData_AllConditionsSummarized.xlsx)


# Precomputed GWAS and PheWAS

The **GWAS** and **PheWAS** tabs in Platform Core Cohorts allow you to visualize precomputed analysis results for phenotypes/diseases and genes, respectively.

{% hint style="info" %}
*These do not reflect the subjects that are part of the cohort that you created.*
{% endhint %}

Platform Core Cohorts currently hosts GWAS and PheWAS analysis results for approximately 150 quantitative phenotypes, such as "LDL direct" and "sitting height", and about 700 diseases.

## Visualize Results from Precomputed Genome-Wide Association Studies (GWAS)

Navigate to the **GWAS** tab and start looking for phenotypes and diseases in the search box. Cohorts will suggest the best matches against any partial input, such as "cancer", you provide. After selecting a phenotype or disease, Cohorts will render a Manhattan plot, by default collapsed to gene level and organized by their respective position in each chromosome.

Circles in the Manhattan plot indicate binary traits, potential associations between genes and diseases. Triangles indicate quantitative phenotypes with regression Beta different from zero, and point up or down to depict positive or negative correlation, respectively.

Hovering over a circle/triangle will display the following information:

* gene symbol
* variant group (see below)
* P-value, both raw and FDR-corrected
* number of carriers of variants of the given type
* number of carriers of variants of any type
* regression Beta

For gene-level results, Cohorts distinguishes five different classes of variants: protein truncating; deleterious; missense; missense with a high ILMN PrimateAI score, indicating likely damaging variants; and synonymous variants. You can limit results to either of these five classes, or select **All** to display all results together.

* Deleterious variants (`del`): the union of all protein-truncating variants (PTVs, defined below), pathogenic missense variants with a PrimateAI score greater than a gene-specific threshold, and variants with a SpliceAI score greater than 0.2.
* Protein-truncating variants (`ptv`): variant consequences matching any of `stop_gained`, `stop_lost`, `frameshift_variant`, `splice_donor_variant`, `splice_acceptor_variant`, `start_lost`, `transcript_ablation`, `transcript_truncation`, `exon_loss_variant`, `gene_fusion`, or `bidirectional_gene_fusion`.
* `missense_all`: all missense variants regardless of their pathogenicity.
* missense, PrimateAI optimized (`missense_pAI_optimized`): only pathogenic missense variants with primateAI score greater than a gene-specific threshold.
* missenses and PTVs (`missenses_and_ptvs_all`): the union of all PTVs, SpliceAI > 0.2 variants and all missense variants regardless of their pathogenicity scores.
* all synonymous variants (`syn`).

To zoom in to a particular chromosome, click the chromosome name underneath the plot, or select the chromosome from the drop down box, which defaults to **Whole genome**.

## Visualize Results from Precomputed Phenome-Wide Association Studies (PheWAS)

To browse PheWAS analysis results by gene, navigate to the **PheWAS** tab and enter a gene of interest into the search box. The resulting Manhattan plot will show phenotypes and diseases organized into a number of categories, such as "Diseases of the nervous system" and "Neoplasms". Click on the name of a category, shown underneath the plot, to display only those phenotypes or diseases, or select a category from the drop down, which defaults to **All**.


# Cohort Analysis

## Cohort Analysis

From the Cohorts menu in the left hand navigation, select a cohort created in **Create Cohort** to begin a cohort analysis.

### Query Details

The query details can be accessed by clicking the triangle next to **Show Query Details**. The query details displays the selections used to create a cohort. The selections can be edited by clicking the **pencil** icon in the top right.

### Charts

1. **Charts** will be open by default. If not, click **Show Charts**.
2. Use the **gear** icon in the top-right to change viewable chart settings.
3. There are four charts available to view summary counts of attributes within a cohort as histogram plots.
4. Click **Hide Charts** to hide the histograms.

### Single Subject Timeline View:

1. Display time-stamped events and observations for a single subject on a timeline.The timeline view is visible to only those subjects which have time-series data.
2. The following attributes are displayed in timeline view:
   * Diagnosed and Self-Reported Diseases:
     * Start and end dates
     * Progression vs. remission
   * Medication and Other Treatments:
     * Prescribed and self-medicated
     * Start date, end date, and dosage at every time point
3. The timeline utilizes age (at diagnosis, at event, at measurement) as the x-axis and attribute name as the y-axis. If the birthdate is not recorded for a subject, then you can switch to Date to visualize data.
4. In the default view, the timeline shows the first five disease data and the first five drug/medication data in the plot. You can choose different attributes or change the order of existing attributes by clicking on the “select attribute” button.
5. The x-axis shows the person’s age in years, with data points initially displayed between ages 0 to 100. You can zoom in by selecting the desired range to visualize data points within the selected age range.
6. Each event is represented by a dot in the corresponding track. Events in the same track can be connected by lines to indicate the start and end period of an event.

{% hint style="info" %}
**Measurement Section**: A summary of measurements (without values) is displayed under the section titled "Measurements and Laboratory Values Available." You can click a link to access the Timeline View for detailed results.

**Drug Section**: The "Drug Name" section lists drug names without repeating the header "Drug Name" for each entry.
{% endhint %}

### Subjects

1. By Default, the **Subjects** tab is displayed.
2. The **Subjects** tab with a list of all subjects matching your criteria is displayed below **Charts** with a link to each Subject by ID and other high-level information. By clicking a subject ID, you will be brought to the data collected at the Subject level.
3. Search for a specific subject by typing the Subject ID into the **Search Subjects** text box.
4. Get all details available on a subject by clicking the hyperlinked Subject ID in the Subject list.

To **Exclude** specific subjects from subsequent analysis, such as marker frequencies or gene-level aggregated views, you can uncheck the box at the beginning of each row in the subject list. You will then be prompted to save any exclusion(s).

You can **Export** the list of subjects either to your Platform Core project's data folder or to your local disk as a TSV file for subsequent use. Any export will omit subjects that you excluded after you saved those changes. For more information, see at the bottom of this page.

#### Remove a Subject

1. Specific subjects can be removed from a Cohort.
2. Select the **Subjects** tab.
3. Subjects in the Cohort, by default are **checked**.
4. To remove a specific subject from a Cohort, **uncheck** the checkbox next to subjects to remove from a Cohort.
5. Check box selections are maintained while browsing through the pages of the subject list.
6. Click **Save Cohort** to save the subjects you would like to exclude.
7. The specific subjects will no longer be counted in all analysis visualizations.
8. The specific excluded subjects will be saved for the Cohort.
9. To add the subjects back to the Cohort, select the checkboxes to **checked** and click **Save Cohort**.

### Structural variant aggregation: Marker Frequency analysis

For each individual cohort, display a table of all observed SVs that overlap with a given gene.

### Marker Frequency

1. Click the **Marker Frequency** tab, then click the **Gene Expression** tab.
2. Down-regulated genes are displayed in blue and Up-regulated genes are displayed in red.
3. A frequency in the Cohort is displayed and the Matching number/Total is also displayed in the chart.
4. Genes can be searched by using the **Search Genes** text box.

### Genes

1. You are brought to the **Gene** tab under the **Gene Summary** sub-tab.
2. Select a Gene by typing the gene name into the **Search Genes** text box.
3. A **Gene Summary** will be displayed that lists information and links to public resources about the selected gene.
4. A cytogenic map will be displayed based on the selected gene and a vertical orange bar represents gene location in the chromosome.
5. Click the **Variants** tab and **Show legend and filters** if it does not open by default.
6. Below the interactive legend, you see a set of analysis tracks: Needle Plot, Primate AI, Pathogenic variants, and Exons.
7. The Needle Plot allows toggling the plot by **gnomAD frequency** and **Sample Count**. Select **Sample Count** in the **Plot by** legend above the plot. You can also filter the plot to only show variants above or below a certain cut-off for gnomAD frequency in percent or absolute sample count.
8. The Needle Plot allows filtering by **PrimateAI** Score.
   * Set a lower (>=) or upper (<=) threshold for the PrimateAI Score to filter variants.
   * Enter the threshold value in the text box located below the gnomadFreq/SampleCount input box.
   * If no threshold value is entered, no filter will be applied.
   * The filter affects both the plot and the table when the “Display only variants shown in the plot above” toggle is enabled.
   * Filter preferences persist across gene views for a seamless experience.
9. The following filters are always shown and can be independently set: **%gnomAD Frequency**, **Sample Count**, **PrimateAI Score**. Changes made to these filters are immediately reflected in both the needle plot and the variant list below.
10. Click on a variant's needle pin to view details about the variant from public resources and counts of variants in the selected cohort by disease category. If you want to view all subjects that carry the given variant, click on the sample count link, which will take you to the list of subjects (see above).
11. Use the Exon zoom bar from each end of the Amino Acid sequence to zoom in on the gene domain to better separate observations.
12. The **Pathogenic Variant** track shows pop-up details with pathogenicity calls, phenotypes, submitter and a link to the ClinVar entry when you hover over the purple triangles.
13. Below the needle plot is a full listing of variants displayed in the needle plot visualization
    * **Display only variants shown in the plot above.** toggle (enabled by default) syncs the table with the Needle Plot. When the toggle is on, the table will display only the variants shown in the Needle Plot, applying all active filters (e.g., variant type, somatic/germline, sample count). When the toggle is off, all reported variants are displayed in the table and table-based filters can be used.
    * **Export to CSV**: When the views are synchronized (toggle on), the filtered list of variants can be exported to a CSV file for further analysis.The `Phenotypes tab` shows a stacked horizontal bar chart which displays molecular breakdown (disease type vs Gene) and subject count for the selected gene.

{% hint style="info" %}
For "Stop Lost" Consequence Variants:

* The **stop\_lost** consequence is mapped as **Frameshift, Stop lost** in the tooltip.
* The **Stop gained|lost** value includes both stop gain and stop loss variants.
* When the Stop gained filter is applied, Stop lost variants will not appear in the plot or table if the "Display only variants shown in the plot above" toggle is enabled
  {% endhint %}

14. The **Gene Expression** tab shows known gene expression data from tissue types in GTEx.
15. The **Genetic Burden Test** will only be available for **de novo** variants only.

## Correlation

{% hint style="info" %}
For every correlation, subjects contained in each count can be viewed by selecting the count on the *bubble* or the count on the X-axis and Y-axis.
{% endhint %}

### Clinical vs. Clinical Attribute Comparison – Bubble Plot

1. Click the **Correlation** tab.
2. In **X-axis category**, select **Clinical**.
3. In **X-axis Attribute**, select a clinical attribute.
4. In **Y-axis category**, select **Clinical**.
5. In **Y-Axis Attribute**, select another clinical attribute.
6. You will be shown a bubble plot comparing the first clinical attribute on the x-axis to second attributes on the y-axis.
7. The size of the bubbles correspond to the number of subjects falling into those categories.

### Molecular vs. Molecular Attribute Comparison – Bubble Plot

To see a breakdown of Somatic Mutations vs. RNA Expression levels perform the following steps:

**This comparison is for a Cancer case.**

1. Click the **Correlation** tab.
2. In **X-axis category**, select **Somatic**.
3. In **X-axis Attribute**, select a gene.
4. In **Y-axis category**, select **RNA expression**.
5. In **Y-Axis Attribute**, type a gene and leave **Reference Type**, **NORMAL**.
6. Click **Continuous** to see violin plots of compared variables.

### Clinical vs. Molecular Attribute Comparison – Bubble Plot

**This comparison is for a Cancer case.**

1. Click the **Correlation** tab.
2. In **X-axis category**, select **Somatic**.
3. In **X-axis Attribute**, type a gene name.
4. In **Y-axis category**, select **Clinical**.
5. In **Y-Axis Attribute**, select a clinical attribute.

### Molecular Breakdown

1. Click the **Molecular Breakdown** tab.
2. In **Enter a clinical Attribute**, select a clinical attribute.
3. In **Enter a gene**, select a gene by typing a gene name.
4. You are shown a stacked bar-chart by the clinical attribute selected values on the Y-axis.
5. For each attribute value the bar represents the % of Subjects with **RNA Expression**, **Somatic Mutation**, and **Multiple Alterations**.

{% hint style="info" %}
For each of the aforementioned bubble plots, you can view the list of subjects by following the link under each subject count associated with an individual bubble or axis label. This will take you to the list of subjects view, see above.
{% endhint %}

### CNV

If there is Copy Number Variant data in the cohort:

1. Click the **CNV** tab.
2. A graph will show CNV a Sample Percentage on the Y-axis and Chromosomes on the X-axis.
3. Any value above Zero is a copy number gain, and any value below Zero is a copy number loss.
4. Click **Chromosome:** to select a specific chromosome position.

### Subject Export for Analysis in Platform Core Bench

Platform Core allows integrated analysis in a computation workspace. You can export your cohort definitions and, in combination with molecular data in your Platform Core project data, perform, for example, a GWAS analysis.

1. Confirm the VCF data for your analysis is in Platform Core project data.
2. From within your Platform Core project, start a Bench workspace. See [Bench](/project/p-bench) for more details.
3. Navigate back to Platform Core Cohorts.
4. Create a Cohort of subjects of interest using [Create a Cohort](/project/p-cohorts/cohorts-create).
5. From the **Subjects** tab click **Export subjects...** from the top-right of the subject list. The file can be downloaded to the browser or Platform Core project data.
6. We suggest using export **...to Data Folder** for immediate access to this data in Bench or other areas of Platform Core.
7. Create another cohort if needed for your Research and complete the last 3 steps.
8. Navigate to the Bench workspace created in the second step.
9. After the workspace has started up, click **Access**.
10. Find the `/Project/` folder in the Workspace file navigation.
11. This folder will contain your cohort files created along with any pipeline output data needed for your workspace analysis.


# Compare Cohorts

You can compare up to four previously created individual cohorts, to view differences in variants and mutations, RNA expression, copy number variation, and distribution of clinical attributes. Once comparisons are created, they are saved in the **Comparisons** left-navigation tab of the Cohorts module.

## Create a comparison view

1. Select **Cohorts** from the left-navigation panel.
2. Select 2 to 4 cohorts already created. If you have not created any cohorts, See *Create a Cohort* documentation.
3. Click **Compare Cohorts** in the right-navigation panel.
4. Note you are now in the **Comparisons** left-navigation tab of the Cohorts module.
5. In the **Charts** section, if the **COHORTS** item is not displayed, click the gear icon in the top right and add **Cohorts** as the first attribute and click **Save**.
6. The **COHORTS** item in the charts panel will provide a count of subjects in each cohort and act as a legend for color representation throughout comparison screens.
7. For each clinical attribute category, a bar chart is displayed. Use the gear icon to select attributes to display in the charts panel.

{% hint style="info" %}
You can share a comparison with other team members in the same Platform Core project. Please refer to the section on "Sharing a Cohort" on "Create a Cohort" for details on sharing, unsharing, deleting, and archiving, which are analogous for sharing comparisons.
{% endhint %}

## Attribute Comparison

1. Select the **Attributes** tab
2. Attribute categories are listed and can be expanded using the down-arrows next to the category names. The categories available are based on cohorts selected. Categories and attributes are part of the Platform Core Cohorts metadata template that map to each Subject.
3. For example, use the drop-down arrow next to **Vital status** to view sub-categories and frequencies across selected cohorts.

## Variants Comparison

1. Select the **Genes** tab
2. Search for a gene of interest using its HUGO/HGNC gene symbol
3. Variants and mutations will be displayed as one needle plot for each cohort that is part of the comparison (see [Genes](/project/p-cohorts/cohorts-analysis#genes) for more details)
4. As additional filter options, you can view only those variants that are occur in every cohort; that are unique to one cohort; that have been observed in at least two cohorts; or any variant.

## Survival Summary

1. Select the **Survival Summary** tab.
2. Attribute categories are listed and can be expanded using the down-arrows next to the category names.
3. Select the drop-down arrow for **Therapeutic interventions**.
4. In each subcategory there is a sum of the subject counts across select cohorts.
5. For each cohort, designated by a color, there is a **Subject count** and **Median survival (years)** column.
6. Type **Malignancy** in the Search Box and an auto-complete dropdown suggests three different attributes.
7. Select **Synchronous malignancy** and the results are automatically opened and highlighted in orange.

## Survival Comparison

1. Click **Survival Comparison** tab.
2. A Kaplan-Meier Curve is rendered based on each cohort.
3. P-Value Displayed at the top of Survival Comparison indicates whether there is statistically significant variance between survival probabilities over time of any pair of cohorts (CI=0.95).

{% hint style="info" %}
When comparing two cohorts, the P-Value is shown above the two survival curves. For three or four cohorts, P-Values are shown as a pair-wise heatmap, comparing each cohort to every other cohort.
{% endhint %}

## Marker Frequency Comparison

1. Select the **Marker Frequency** tab.
2. Select either **Gene expression** (default), **Somatic mutation**, or **Copy number variation**
3. For gene expression (up- versus down-regulated) and for copy number variation (gain versus loss), Cohorts will display a list of all genes with bidirectional barcharts
4. For somatic mutations, the barcharts are unidirectional and indicate the percentage of samples with a mutation in each gene per cohort.
5. Bars are color-coded by cohort, see the accompanying legend.
6. Each row shows P-value(s) resulting from pairwise comparison of all cohorts. In the case of comparing two cohorts, the numerical P-value will be displayed in the table. In the case of comparing three or more cohorts, the pairwise P-values are shown as a triangular heatmap, with details available as a tooltip.

## Correlation Comparison

1. Select the **Correlation** tab.
2. Similar to the single-cohort view (**Cohort Analysis | Correlation**), choose two clinical attributes and/or genes to compare.
3. Depending on the available data types for the two selections (categorical and/or continuous), Cohorts will display a bubble plot, violin plot, or scatter plot.


# Cohorts Data in Base

Platform Core Cohorts data can be viewed in a Platform Core Project Base instance as a *shared database*. A shared database in Platform Core Base operates as a database view. To use this feature, enable Base for your project prior to starting any Platform Core Cohorts ingestions. See [Base](/project/p-base) for more information on enabling this feature in your Platform Core Project.

## Platform Core Cohorts Base Tables

After ingesting data into your project, select Phenotypic and Molecular data are available to view in Base. See Cohorts [Import](/project/p-cohorts/cohorts-import) for instruction on importing data sets into Cohorts.

1. Post ingestion, data will be represented in Base.
2. Select **BASE** from the Platform Core left navigation and click **Query**.
3. Under the New Query window, a list of tables is displayed. Expand the **Shared Database for Project \<your\_project>**.
4. Cohorts tables will be displayed.
5. To preview the table and fields click each listed view.
6. Clicking any of these views then selecting **Preview** on the right-hand side will show you a preview of the data in the tables.

{% hint style="info" %}
If your ingestion includes Somatic variants, there will be two molecular tables: *ANNOTATED\_SOMATIC\_MUTATIONS* and *ANNOTATED\_VARIANTS*. All ingestions will include a *PHENOTYPE* table.
{% endhint %}

{% hint style="info" %}
The *PHENOTYPE* table includes a harmonized set that is collected across all data ingestions and is not representative of all data ingested for the Subject or Sample. Sample information is also displayed in this table, if applicable. Sample information drives the annotation process if molecular data is included in the ingestion. That data is stored in the *PHENOTYPE* table.
{% endhint %}

## Phenotype Data

| Field Name             | Type    | Description                                                |
| ---------------------- | ------- | ---------------------------------------------------------- |
| SAMPLE\_BARCODE        | STRING  | Sample Identifier                                          |
| SUBJECTID              | STRING  | Identifer for Subject entity                               |
| STUDY                  | STRING  | Study designation                                          |
| AGE                    | NUMERIC | Age in years                                               |
| SEX                    | STRING  | Sex field to drive annotation                              |
| POPULATION             | STRING  | Population Designation for 1000 Genomes Project            |
| SUPERPOPULATION        | STRING  | Superpopulation Designation from 1000 Genomes Project      |
| RACE                   | STRING  | Race according to NIH standard                             |
| CONDITION\_ONTOLOGIES  | VARIANT | Diagnosis Ontology Source                                  |
| CONDITION\_IDS         | VARIANT | Diagnosis Concept Ids                                      |
| CONDITIONS             | VARIANT | Diagnosis Names                                            |
| HARMONIZED\_CONDITIONS | VARIANT | Diagnosis High-level concept to drive UI                   |
| LIBRARYTYPE            | STRING  | Seqencing technology                                       |
| ANALYTE                | STRING  | Substance sequenced                                        |
| TISSUE                 | STRING  | Tissue source                                              |
| TUMOR\_OR\_NORMAL      | STRING  | Tumor designation for somatic                              |
| GENOMEBUILD            | STRING  | Genome Build to drive annotations - hg38 only              |
| SAMPLE\_BARCODE\_VCF   | STRING  | Sample ID from VCF                                         |
| AFFECTED\_STATUS       | NUMERIC | Affected, Unaffected, or Unknown for Family Based Analysis |
| FAMILY\_RELATIONSHIP   | STRING  | Relationship designation for Family Based Analysis         |

## Sample Information

| Field Name      | Type   | Description                                |
| --------------- | ------ | ------------------------------------------ |
| SAMPLE\_BARCODE | STRING | Original sample barcode used in VCF column |
| SUBJECTID       | STRING | Original identifier for the subject record |
| DATATYPE        | ARRAY  | The categorization of molecular data       |
| TECHNOLOGY      | ARRAY  | The sequencing method                      |
| CREATEDATE      | DATE   | Date and time of record creation           |
| LASTUPDATEDATE  | DATE   | Date and time of last update of record     |

## Sample Attribute

This table is an entity-attribute value table of supplied sample data matching Cohorts accepted attributes.

| Field Name       | Type    | Description                                |
| ---------------- | ------- | ------------------------------------------ |
| SAMPLE\_ BARCODE | STRING  | Original sample barcode used in VCF column |
| SUBJECTID        | STRING  | Original identifier for the subject record |
| ATTRIBUTE\_NAME  | STRING  | Cohorts meta-data driven field name        |
| ATTRIBUTE\_VALUE | VARIANT | List of values entered for the field       |

## Study Information

| Field Name     | Type   | Description                     |
| -------------- | ------ | ------------------------------- |
| NAME           | STRING | Study name                      |
| CREATEDATE     | DATE   | Date and time of study creation |
| LASTUPDATEDATE | DATE   | Data and time of record update  |

## Subject

| Field          | Type   | Description                                 |
| -------------- | ------ | ------------------------------------------- |
| SUBJECTID      | STRING | Original identifier for the subject record  |
| AGE            | FLOAT  | Age entered on subject record if applicable |
| SEX            | STRING | -                                           |
| ETHNICITY      | STRING | -                                           |
| STUDY          | STRING | Study subject belongs to                    |
| CREATEDATE     | DATE   | Date and time of record creation            |
| LASTUPDATEDATE | DATE   | Date and time of record update              |

## Subject Attribute

This table is an entity-attribute value table of supplied subject data matching Cohorts accepted attributes.

| Field            | Type    | Description                                |
| ---------------- | ------- | ------------------------------------------ |
| SUBJECTID        | STRING  | Original identifier for the subject record |
| ATTRIBUTE\_NAME  | STRING  | Cohorts meta-data driven field name        |
| ATTRIBUTE\_VALUE | VARIANT | List of values entered for the field       |

## Disease

<table><thead><tr><th>Field</th><th width="208">Type</th><th>Description</th></tr></thead><tbody><tr><td>SUBJECTID</td><td>STRING</td><td>Original identifier for the subject record</td></tr><tr><td>TERM</td><td>STRING</td><td>Code for disease term</td></tr><tr><td>OCCURRENCES</td><td>STRING</td><td>List of occurrence related data</td></tr></tbody></table>

## Drug Exposure

| Field       | Type   | Description                                      |
| ----------- | ------ | ------------------------------------------------ |
| SUBJECTID   | STRING | Original identifier for the subject record       |
| TERM        | STRING | Code for drug term                               |
| OCCURRENCES | STRING | List of occurrence related data of drug exposure |

## Measurement

| Field       | Type   | Description                                                       |
| ----------- | ------ | ----------------------------------------------------------------- |
| SUBJECTID   | STRING | Original identifier for the subject record                        |
| TERM        | STRING | Code for measurement term                                         |
| OCCURRENCES | STRING | List of occurrences and values related to lab or measurement data |

## Procedure

| Field       | Type   | Description                                           |
| ----------- | ------ | ----------------------------------------------------- |
| SUBJECTID   | STRING | Original identifier for the subject record            |
| TERM        | STRING | Code for procedure term                               |
| OCCURRENCES | STRING | List of occurrences and values related procedure data |

## Annotated Variants

This table will be available for all projects with ingested molecular data

| **Field Name**  | **Type** | **Description**                                                                |
| --------------- | -------- | ------------------------------------------------------------------------------ |
| SAMPLE\_BARCODE | STRING   | Original sample barcode used in VCF column                                     |
| STUDY           | STRING   | Study designation                                                              |
| GENOMEBUILD     | STRING   | Only hg38 is supported                                                         |
| CHROMOSOME      | STRING   | Chromosome without 'chr' prefix                                                |
| CHROMOSOMEID    | NUMERIC  | Chromosome ID: 1..22, 23=X, 24=Y, 25=Mt                                        |
| DBSNP           | STRING   | dbSNP Identifiers                                                              |
| VARIANT\_KEY    | STRING   | Variant ID in the form "1:12345678:12345678:C"                                 |
| NIRVANA\_VID    | STRING   | Broad Institute VID: "1-12345678-A-C"                                          |
| VARIANT\_TYPE   | STRING   | Description of Variant Type (e.g. SNV, Deletion, Insertion)                    |
| VARIANT\_CALL   | NUMERIC  | 1=germline, 2=somatic                                                          |
| DENOVO          | BOOLEAN  | true / false                                                                   |
| GENOTYPE        | STRING   | "G\|T"                                                                         |
| READ\_DEPTH     | NUMERIC  | Sequencing read depth                                                          |
| ALLELE\_COUNT   | NUMERIC  | Counts of each alternate allele for each site across all samples               |
| ALLELE\_DEPTH   | STRING   | Unfiltered count of reads that support a given allele for an individual sample |
| FILTERS         | STRING   | Filter field from VCF. If all filters pass, field is PASS                      |
| ZYGOSITY        | NUMERIC  | 0 = hom ref, 1 = het ref/alt, 2 = hom alt, 4 = hemi alt                        |
| GENEMODEL       | NUMERIC  | 1=Ensembl, 2=RefSeq                                                            |
| GENE\_HGNC      | STRING   | HUGO/HGNC gene symbol                                                          |
| GENE\_ID        | STRING   | Ensembl gene ID ("ENSG00001234")                                               |
| GID             | NUMERIC  | NCBI Entrez Gene ID (RefSeq) or numerical part of Ensembl ENSG ID              |
| TRANSCRIPT\_ID  | STRING   | Ensembl ENST or RefSeq NM\_                                                    |
| CANONICAL       | STRING   | Transcript designated 'canonical' by source                                    |
| CONSEQUENCE     | STRING   | missense, stop gained, intronic, etc.                                          |
| HGVSC           | STRING   | The HGVS coding sequence name                                                  |
| HGVSP           | STRING   | The HGVS protein sequence name                                                 |

## Annotated Somatic Mutations

This table will only be available for data sets with ingested *Somatic* molecular data.

| **Field Name**  | **Type** | **Description**                                                                            |
| --------------- | -------- | ------------------------------------------------------------------------------------------ |
| SAMPLE\_BARCODE | STRING   | Original sample barcode, used in VCF column                                                |
| SUBJECTID       | STRING   | Identifier for Subject entity                                                              |
| STUDY           | STRING   | Study designation                                                                          |
| GENOMEBUILD     | STRING   | Only hg38 is supported                                                                     |
| CHROMOSOME      | STRING   | Chromosome without 'chr' prefix                                                            |
| DBSNP           | NUMERIC  | dbSNP Identifiers                                                                          |
| VARIANT\_KEY    | STRING   | Variant ID in the form "1:12345678:12345678:C"                                             |
| MUTATION\_TYPE  | NUMERIC  | Rank of consequences by expected impact: 0 = Protein Truncating to 40 = Intergenic Variant |
| VARIANT\_CALL   | NUMERIC  | 1=germline, 2=somatic                                                                      |
| GENOTYPE        | STRING   | "G\|T"                                                                                     |
| REF\_ALLELE     | STRING   | Reference allele                                                                           |
| ALLELE1         | STRING   | First allele call in the tumor sample                                                      |
| ALLELE2         | STRING   | Second allele call in the tumor sample                                                     |
| GENEMODEL       | NUMERIC  | 1=Ensembl, 2=RefSeq                                                                        |
| GENE\_HGNC      | STRING   | HUGO/HGNC gene symbol                                                                      |
| GENE\_ID        | STRING   | Ensembl gene ID ("ENSG00001234")                                                           |
| TRANSCRIPT\_ID  | STRING   | Ensembl ENST or RefSeq NM\_                                                                |
| CANONICAL       | BOOLEAN  | Transcript designated 'canonical' by source                                                |
| CONSEQUENCE     | STRING   | missense, stop gained, intronic, etc.                                                      |
| HGVSP           | STRING   | HGVS nomenclature for AA change: p.Pro72Ala                                                |

## Annotated Copy Number Variants

This table will only be available for data sets with ingested *CNV* molecular data.

| **Field Name**       | **Type** | **Description**                                                                     |
| -------------------- | -------- | ----------------------------------------------------------------------------------- |
| SAMPLE\_BARCODE      | STRING   | Sample barcode used in the original VCF                                             |
| GENOMEBUILD          | STRING   | Genome build, always 'hg38'                                                         |
| NIRVANA\_VID         | STRING   | Variant ID of the form 'chr-pos-ref-alt'                                            |
| CHRID                | STRING   | Chromosome without 'chr' prefix                                                     |
| CID                  | NUMERIC  | Numerical representation of the chromosome, X=23, Y=24, Mt=25                       |
| GENE\_ID             | STRING   | NCBI or Ensembl gene identifier                                                     |
| GID                  | NUMERIC  | Numerical part of the gene ID; for Ensembl, we remove the 'ENSG000..' prefix        |
| START\_POS           | NUMERIC  | First affected position on the chromosome                                           |
| STOP\_POS            | NUMERIC  | Last affected position on the chromosome                                            |
| VARIANT\_TYPE        | NUMERIC  | 1 = copy number gain, -1 = copy number loss                                         |
| COPY\_NUMBER         | NUMERIC  | Observed copy number                                                                |
| COPY\_NUMBER\_CHANGE | NUMERIC  | Fold-chang of copy number, assuming 2 for diploid and 1 for haploid as the baseline |
| SEGMENT\_VALUE       | NUMERIC  | Average FC for the identified chromosomal segment                                   |
| PROBE\_COUNT         | NUMERIC  | Probes confirming the CNV (arrays only)                                             |
| REFERENCE            | NUMERIC  | Baseline taken from normal samples (1) or averaged disease tissue (2)               |
| GENE\_HGNC           | STRING   | HUGO/HGNC gene symbol                                                               |

## Annotated Structural Variants

This table will only be available for data sets with ingested *SV* molecular data. Note that Platform Core Cohorts stores copy number variants in a separate table.

| **Field Name**       | **Type**  | **Description**                                                                                                                                                           |
| -------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SAMPLE\_BARCODE      | STRING    | Sample barcode used in the original VCF                                                                                                                                   |
| GENOMEBUILD          | STRING    | Genome build, always 'hg38'                                                                                                                                               |
| NIRVANA\_VID         | STRING    | Variant ID of the form 'chr-pos-ref-alt'                                                                                                                                  |
| CHRID                | STRING    | Chromosome without 'chr' prefix                                                                                                                                           |
| CID                  | NUMERIC   | Numerical representation of the chromosome, X=23, Y=24, Mt=25                                                                                                             |
| BEGIN                | NUMERIC   | First affected position on the chromosome                                                                                                                                 |
| END                  | NUMERIC   | Last affected position on the chromosome                                                                                                                                  |
| BAND                 | STRING    | Chromosomal band                                                                                                                                                          |
| QUALIITY             | NUMERIC   | Quality from the original VCF                                                                                                                                             |
| FILTERS              | ARRAY     | Filters from the original VCF                                                                                                                                             |
| VARIANT\_TYPE        | STRING    | Insertion, deletion, indel, tandem\_duplication, translocation\_breakend, inversion ("INV"), short tandem repeat ("STR2")                                                 |
| VARIANT\_TYPE\_ID    | NUMERIC   | 21=insertion, 22=deletion, 23=indel, 24=tandem\_duplication, 25=translocation\_breakend, 26=inversion ("INV"), 27=short tandem repeat ("STR2")                            |
| CIPOS                | ARRAY     | Confidence interval around first position                                                                                                                                 |
| CIEND                | ARRAY     | Confidence interval around last position                                                                                                                                  |
| SVLENGTH             | NUMERIC   | Overall size of the structural variant                                                                                                                                    |
| BONDCHR              | STRING    | For translocations, the other affected chromosome                                                                                                                         |
| BONDCID              | NUMERIC   | For translocations, the other affected chromosome as a numeric value, X=23, Y=24, Mt=25                                                                                   |
| BONDPOS              | STRING    | For translocations, positions on the other affected chromosome                                                                                                            |
| BONDORDER            | NUMERIC   | 3 or 5: Whether this fragment (the current variant/VID) "receives" the other chromosome's fragment on it's 3' end, or attaches to the 5' of the other chromosome fragment |
| GENOTYPE             | STRING    | Called genotype from the VCF                                                                                                                                              |
| GENOTYPE\_QUALITY    | NUMERIC   | Genotype call quality                                                                                                                                                     |
| READCOUNTSSPLIT      | ARRAY     | Read counts                                                                                                                                                               |
| READCOUNTSPAIRED     | ARRAY     | Read counts, paired end                                                                                                                                                   |
| REGULATORYREGIONID   | STRING    | Ensembl ID for the affected regulatory region                                                                                                                             |
| REGULATORYREGIONTYPE | STRING    | Type of the regulatory region                                                                                                                                             |
| CONSEQUENCE          | ARRAY     | Variant consequence according to SequenceOntology                                                                                                                         |
| TRANSCRIPTID         | STRING    | Ensembl of RefSeq transcript identifier                                                                                                                                   |
| TRANSCRIPTBIOTYPE    | STRING    | Biotype of the transcript                                                                                                                                                 |
| INTRONS              | STRING    | Count of impacted introns out of the total number of introns, specified as "M/N"                                                                                          |
| GENEID               | STRING    | Ensembl or RefSeq gene identifier                                                                                                                                         |
| GENEHGNC             | STRING    | HUGO/HGNC gene symbol                                                                                                                                                     |
| ISCANONICAL          | BOOLEAN   | Is the transcript ID the canonical one according to Ensembl?                                                                                                              |
| PROTEINID            | STRING    | RefSeq or Ensembl protein ID                                                                                                                                              |
| SOURCEID             | NUMERICAL | Gene model: 1=Ensembl, 2=RefSeq                                                                                                                                           |

## Raw RNAseq data tables for genes and transcripts

These tables will only be available for data sets with ingested *RNAseq* molecular data.

Table for gene quantification results:

| **Field Name**    | **Type**  | **Description**                                                                   |
| ----------------- | --------- | --------------------------------------------------------------------------------- |
| GENOMEBUILD       | STRING    | Genome build, always 'hg38'                                                       |
| STUDY\_NAME       | STRING    | Study designation                                                                 |
| SAMPLE\_BARCODE   | STRING    | Sample barcode used in the original VCF                                           |
| LABEL             | STRING    | Group label specified during import: Case or Control, Tumor or Normal, etc.       |
| GENE\_ID          | STRING    | Ensembl or RefSeq gene identifier                                                 |
| GID               | NUMERIC   | Numerical part of the gene ID; for Ensembl, we remove the 'ENSG000..' prefix      |
| GENE\_HGNC        | STRING    | HUGO/HGNC gene symbol                                                             |
| SOURCE            | STRING    | Gene model: 1=Ensembl, 2=RefSeq                                                   |
| TPM               | NUMERICAL | Transcripts per million                                                           |
| LENGTH            | NUMERICAL | The length of the gene in base pairs.                                             |
| EFFECTIVE\_LENGTH | NUMERICAL | The length as accessible to RNA-seq, accounting for insert-size and edge effects. |
| NUM\_READS        | NUMERICAL | The estimated number of reads from the gene. The values are not normalized.       |

The corresponding transcript table uses TRANSCRIPT\_ID instead of GENE\_ID and GENE\_HGNC.

## Differential expression tables for genes and transcripts

These tables will only be available for data sets with ingested *RNAseq* molecular data.

Table for differential gene expression results:

| **Field Name**       | **Type**  | **Description**                                                              |
| -------------------- | --------- | ---------------------------------------------------------------------------- |
| GENOMEBUILD          | STRING    | Genome build, always 'hg38'                                                  |
| STUDY\_NAME          | STRING    | Study designation                                                            |
| SAMPLE\_BARCODE      | STRING    | Sample barcode used in the original VCF                                      |
| CASE\_LABEL          | STRING    | Study designation                                                            |
| GENE\_ID             | STRING    | Ensembl or RefSeq gene identifier                                            |
| GID                  | NUMERIC   | Numerical part of the gene ID; for Ensembl, we remove the 'ENSG000..' prefix |
| GENE\_HGNC           | STRING    | HUGO/HGNC gene symbol                                                        |
| SOURCE               | STRING    | Gene model: 1=Ensembl, 2=RefSeq                                              |
| BASEMEAN             | NUMERICAL |                                                                              |
| FC                   | NUMERICAL | Fold-change                                                                  |
| LFC                  | NUMERICAL | Log of the fold-change                                                       |
| LFCSE                | NUMERICAL | Standard error for log fold-change                                           |
| PVALUE               | NUMERICAL | P-value                                                                      |
| CONTROL\_SAMPLECOUNT | NUMERICAL | Number of samples used as control                                            |
| CONTROL\_LABEL       | NUMERICAL | Label used for controls                                                      |

The corresponding transcript table uses TRANSCRIPT\_ID instead of GENE\_ID and GENE\_HGNC.


# Oncology Walk-through

This walk-through is intended to represent a typical workflow when building and studying a cohort of oncology cases.

{% embed url="<https://www.youtube.com/watch?v=wr19Y8BVaQ4&list=PLKRu7cmBQlaiQT6Giou9aSkZ4C0LMIGbc&index=8>" %}
Multi-Omic Cancer Workflow
{% endembed %}

## Create a Cancer Cohort and View Subject Details

1. Click **Create Cohort** button.
2. Select the following studies to add to your cohort:
   1. TCGA – BRCA – Breast Invasive Carcinoma
   2. TCGA – Ovarian Serous Cystadenocarcinoma
3. Add a **Cohort Name** = TCGA Breast and Ovarian\_1472
4. Click on **Apply**.
5. Expand **Show query details** to see the study makeup of your cohort.
6. **Charts** will be open by default. If not, click **Show charts**
7. Use the gear icon in the top-right to change viewable chart settings.

{% hint style="info" %}
**Disease Type**, **Histological Diagnosis**, **Technology**, **Overall Survival** have interesting data about this cohort.
{% endhint %}

8. The **Subject** tab with all Subjects list is displayed below Charts with a link to each Subject by ID and other high-level information, like Data Types measured and reported. By clicking a subject ID, you will be brought to the data collected at the Subject level.
9. Search for subject **TCGA-E2-A14Y** and view the data about this Subject.
10. Click the **TCGA-E2-A14Y** Subject ID link to view clinical data for this Subject that was imported via the metadata.tsv file on ingest.

{% hint style="info" %}
The subject is a 35 year old Female with vital status and other phenotypes that feed up into the **Subject** attribute selection criteria when creating or editing cohorts.
{% endhint %}

11. Click **X** to close the Subject details.
12. Click **Hide charts** to increase interactive landscape.

## Data Analysis, Multi-Omic Biomarker Discovery, and Interpretation

1. Click the **Marker Frequency** tab, then click the **Somatic Mutation** tab.
2. Review the gene list and mutation frequencies.
3. You will notice PIK3CA has a high rate of mutation in the Cohort (ranked 2nd with 33% mutation frequency in 326 of the 987 Subjects that have Somatic Mutation data in this cohort).
   * We can investigate if subjects with PIK3CA mutations have changes in PIK3CA RNA Expression?
4. Click the **Gene Expression** tab, search for **PIK3CA**
   * PIK3CA RNA is down-regulated in 27% of the subjects relative to normal samples.
     1. Switch from **normal** to **disease** Reference where the Subject’s denominator is the median of all disease samples in your cohort.
     2. The count of matching vs. total subjects that have PIK3CA up-regulated RNA which may indicate a distinctive sub-phenotype.
5. Click on the **PIK3CA** gene link in the **Gene Expression** table.
6. You are shown the **Gene** tab under the **Gene Summary** sub-tab that lists information and links to public resources about PIK3CA.
7. Click the **Variants** tab and **Show legend and filters** if it does not open by default.
8. Below the interactive legend you see a set of analysis tracks: Needle Plot, Primate AI, Pathogenic variants, and Exons.
9. The Needle Plot allows toggling the plot by **gnomAD frequency** and **Sample Count**. Select **Sample Count** in the **Plot by** legend above the plot.
   * There are 87 mutations distributed across the 1068 amino acid sequence, listed below the analysis tracks. These can be exported via the icon into a table.
10. We know that missense variants can severely disrupt translated protein activity. Deselect all **Variant Types** except for **Missense** from the **Show Variant Type** legend above the needle plot.
    * Many mutations are in the functional domains of the protein as seen by the colored boxes and labels on the x-axis of the Needle Plot.
11. Hover over the variant with the highest sample count in the yellow **PI3Ka** protein domain.
    * The pop-up shows variant details for the 64 Subjects observed with it: 63 in the Breast Cancer study and 1 in the Ovarian Cancer Study.
12. Use the Exon zoom bar from each end of the Amino Acid sequence to zoom in to the **PI3Ka** domain to better separate observations.
13. There are three different missense mutations at this locus changing the wildtype Glutamine at different frequencies to Lysine (64), Glycine (6), or Alanine (2).
14. The **Pathogenic Variant** Track shows 7 ClinVar entries for mutations stacked at this locus affecting amino acid 545. Pop up details with pathogenicity calls, phenotypes, submitter and a link to the ClinVar entry is seen by hovering over the purple triangles.
15. Look at the **Primate AI** track and high Primate AI score.
    * **Primate AI** track displays scores for potential missense variants, based on polymorphisms observed in primate species. Points above the dashed line for the 75th percentile may be considered likely pathogenic as cross-species sequence is highly conserved; you often see high conservancy at the functional domains. Points below the 25th percentile may be considered "likely benign".
16. Click the **Expression** tab and notice that normal Breast and normal Ovarian tissue have relatively high PIK3CA RNA Expression in GTex RNAseq tissue data but ubiquitously expressed.


# Rare Genetic Disorders Walk-through

## Cohorts Walk-through: Rare Genetic Disorders

This walk-through is meant to represent a typical workflow when building and studying a cohort of rare genetic disorder cases.

## Login and Create a new Platform Core Project

Create a new Project to track your study:

1. Log in to Platform Core
2. Navigate to **Projects**.
3. Create a new project using the **+Create** button.
4. Give your project a name and region and click **Save**.
5. Navigate to the Platform Core Cohorts module by selecting **Projects > your\_project > Cohorts > Cohorts**.

## Create and Review a Rare Disease Cohort

1. Navigate to **Projects > your\_project > Cohorts > Cohorts**.
2. Click **Create Cohort** button.
3. Enter a name for your cohort, for example "Rare Disease + 1kGP" at the top, left of the pencil icon.
4. From the Public Data Sets list select:
   * DRAGEN-1kGP
   * All Rare genetic disease cohorts

{% hint style="info" %}
A cohort can also be created based on Technology, Disease Type and Tissue.
{% endhint %}

5. Under Selected Conditions in right panel, click on **Apply**
6. A new page opens with your cohort in a top-level tab.
7. Expand **Query Details** to see the study makeup of your cohort.
8. A set of 4 Charts will be open by default. If they are not, click **Show Charts**.
   * Use the gear icon in the top-right of the Charts pane to change chart settings.
9. The bottom section is demarcated by 8 tabs (Subjects, Marker Frequency, Genes, GWAS, PheWAS, Correlation, Molecular Breakdown, CNV).
10. The **Subjects** tab displays a list of exportable Subject IDs and attributes.
    * Clicking on a **Subject ID** link opens up a Subject details page.

## Analyze Your Rare Disease Cohort Data

1. A recent GWAS publication identified 10 risk genes for intellectual disability (ID) and autism. Our task is to evaluate them in Platform Core Cohorts: TTN, PKHD1, ANKRD11, ARID1B, ASXL3, SCN2A, FHL1, KMT2A, DDX3X, SYNGAP1.
2. First we **Hide charts** for more visual space.
3. Click the **Genes** tab where you need to query a gene to see and interact with results.
4. Type **SCN2A** into the Gene search field and select it from autocomplete dropdown options.
5. The **Gene Summary** tab now lists information and links to public resources about SCN2A.
6. Click on the **Variants** tab to see an interactive Legend and analysis tracks.
   * The Needle Plot displays **gnomAD Allele Frequency** for variants in your cohort. Some are in SCN2A conserved protein domains.
   * In Legend, switch the **Plot by** option to **Sample Count** in your cohort.
   * In Legend, uncheck all **Variant Types** except **Stop gained**. Now you should see 7 variants.
   * Hover over pin heads to see pop-up information about particular variants.
7. The **Primate AI** track displays Scores for potential missense variants, based on polymorphisms observed in primate species. Points above the dashed line for the 75th percentile may be considered "likely pathogenic" as cross-species sequence is highly conserved; you often see high conservancy at the functional domains. Points below the 25th percentile may be considered "likely benign".
8. The **Pathogenic variants** track displays markers from ClinVar color-coded by variant type. Hover over them to see pop-ups with more information.
9. The **Exons** track shows mRNA exon boundaries with click and zoom functionality at the ends.
10. Below the Needle Plot and analysis tracks is a list of Variants observed in the selected cohort.
    * The **Export Gene Variants** table icon is above the legend on right side.
11. Click on the **Gene Expression** tab to see a Bar chart of 50 normal tissues from GTEx in transcripts per million (TPM). SCN2A is highly expressed in certain brain tissues, indicating specificity to where good markers for intellectual disability and autism could be expected.
12. As a final exercise in discovering good markers, click on the tab for **Genetic Burden Test**. The table here associates **Phenotypes** with **Mutations Observed** in each Study selected for our cohort, alongside **Mutations Expected** to derive p-values. Given all considerations above, SCN2A is good marker for intellectual disability (p < 1.465 x 10 -22) and autism (p < 5.290 x 10 -9).
13. Continue to check the other genes of interest in step 1.


# Public Data Sets

Platform Core Cohorts comes front-loaded with a variety of publicly accessible data sets, covering multiple disease areas and also including healthy individuals.

| Data set             | Samples                                           | Diseases/Phenotypes                                                                                     | Reference                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| -------------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1kGP-DRAGEN          | 3202 WGS: 2504 original samples plus 698 relateds | Presumed healthy                                                                                        | [DRAGEN reanalysis of the 1000 Genomes Dataset](https://aws.amazon.com/blogs/industries/dragen-reanalysis-of-the-1000-genomes-dataset-now-available-on-the-registry-of-open-data/)                                                                                                                                                                                                                                                                                                                                                                                      |
| DDD                  | 4293 (3664 affected), *de novos* only             | Developmental disorders                                                                                 | [McRae et al., Nature 19:1194-1196](https://www.nature.com/articles/nature21062)                                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| EPI4K                | 356, *de novos* only                              | Epilepsy                                                                                                | [Epi4K Consortium, Nature 501:217-221](https://www.nature.com/articles/nature12439)                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| ASD Cohorts          | 6786 (4266 affected), *de novos* only             | Autism Spectrum disorder                                                                                | <p><a href="https://doi.org/10.1016/j.neuron.2012.04.009">Iossifov et al. Neuron 74:285-299</a>;<br><a href="https://doi.org/10.1038/nature13908">Iossifov et al. Nature 498:216-221</a>;<br><a href="https://doi.org/10.1038/nature10989">O'Roak et al. Nature 485:246-250</a>;<br><a href="https://doi.org/10.1038/nature10945">Sanders et al. Nature 485:237-241</a>;<br><a href="https://doi.org/10.1016/j.neuron.2015.09.016">Sanders et al. Neuron 87:1215-1233</a>;<br><a href="https://doi.org/10.1038/nature13772">De Rubeis et al. Nature 515:209-215</a></p> |
| De Ligt *et al.*     | 100, *de novos* only                              | Intellectual disability                                                                                 | [De Ligt et al., N Engl J Med 367:1921-1929](https://www.nejm.org/doi/full/10.1056/NEJMoa1206524)                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| Homsy *et al.*       | 1213, *de novos* only                             | Congenital heart disease (HP:0030680)                                                                   | [Homsy et al., Science 350:1262-1266](https://www.science.org/doi/10.1126/science.aac9396)                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| Lelieveld *et al.*   | 820, *de novos* only                              | Intellectual disability                                                                                 | [Lelieveld et al., Nature Neuroscience19:1194-1196](https://www.nature.com/articles/nn.4352)                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| Rauch *et al.*       | 51, *de novos* only                               | Intellectual disability                                                                                 | [Rauch et al., Lancet 380:1674-1682](https://www.sciencedirect.com/science/article/pii/S0140673612614809)                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| Rare Genomes Project | 315 WES (112 pedigrees)                           | Various                                                                                                 | <https://raregenomes.org/>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| TCGA                 | ca. 4200 WES, ca. 4000 RNAseq                     | 12 tumor types                                                                                          | <https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| GEO                  | RNAseq                                            | Auto-immune disorders, incl. asthma, arthritis, SLE, MS, Crohn's disease, Psoriasis, Sjögren's Syndrome | For GEO/GSE study identifiers, please refer to the in-product list of studies                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
|                      | RNAseq                                            | Kidney diseases                                                                                         | For GEO/GSE study identifiers, please refer to the in-product list of studies                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
|                      | RNAseq                                            | Central nervous system diseases                                                                         | For GEO/GSE study identifiers, please refer to the in-product list of studies                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
|                      | RNAseq                                            | Parkinson's disease                                                                                     | For GEO/GSE study identifiers, please refer to the in-product list of studies                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |


# Details

The project details page contains the properties of the project, such as the location, owner, storage and linked bundles. This is also the place where you add assets in the form of linked [bundles](https://help.ica.illumina.com/home/h-bundles).

The project details are configured during project creation and may be updated by the project owner, entities with the project Adminstrator role, and tenant administrators.

## Adding Linked Bundle Assets

1. Click the **Edit** button at the top of the Details page.
2. Click the **+** button, under LINKED BUNDLES.
3. Click on the desired bundle, then click the **Link** button.
4. Click **Save**.

If your linked bundle contained a pipeline, then it will appear in **Projects > your\_project > Flow > Pipelines**.

## Details List

<table><thead><tr><th width="230">Detail</th><th>Description</th></tr></thead><tbody><tr><td>Name</td><td>Name of the project unique within the tenant. Alphanumerics, underscores, dashes, and spaces are permitted.</td></tr><tr><td>Short Description</td><td>Short description of the project</td></tr><tr><td>Project Owner</td><td>Owner of the project (has Administrator access to the project)</td></tr><tr><td>Storage Configuration</td><td>Storage configuration to use for data stored in the project</td></tr><tr><td>User Tags</td><td>User tags on the project</td></tr><tr><td>Technical Tags</td><td>Technical tags on the project</td></tr><tr><td>Metadata Model</td><td>Metadata model assigned to the project</td></tr><tr><td>Project Location</td><td>Project region where data is stored and pipelines are executed. Options are derived from the Entitlement(s) assigned to user account, based on the purchased subscription</td></tr><tr><td>Storage Bundle</td><td>Storage bundle assigned to the project. Derived from the selected Project Location based on the Entitlement in the purchased subscription</td></tr><tr><td>Billing Mode</td><td>Billing mode assigned to the project</td></tr><tr><td>Data sharing</td><td>Enables data and samples in the project to be linked to other projects</td></tr></tbody></table>

## Billing Mode

A project's billing mode determines the strategy for how costs are charged to billable accounts.

<table><thead><tr><th width="133">Billing Mode</th><th>Description</th></tr></thead><tbody><tr><td>Project</td><td>All incurred costs will be charged to the tenant of the project owner</td></tr><tr><td>Tenant</td><td>Incurred costs will be charged to the tenant of the user owning the project resource (ie, data, analysis). The only exceptions are base tables and queries, as well as bench compute and storage costs, which are always billed to the project owner.</td></tr></tbody></table>

For example, with billing mode set to **Tenant**, if tenant A has created a project resource and uses it in their project, then tenant A will pay for the resource data, compute costs and storage costs of any output they generate within the project. When they share the project with tenant B, then tenant B will pay the compute and storage for the data which they generate in that project. Put simply, in billing mode tenant, the person who generates data pays for the processing and storage of that data, regardless of who owns the actual project.

{% hint style="info" %}
[**Bench**](/project/p-bench) **workspaces always use project billing** even when tenant billing is selected on their project.
{% endhint %}

If the project billing mode is updated after the project has been created, the updated billing mode will only be applied to resources generated after the change.

If you are using your own S3 storage, then the billing mode impacts where collaborator data is stored.

* Project billing will result in using your S3 storage for the data.
* Tenant billing will result in collaborator data being stored in Illumina-managed storage instead of your own S3 storage.
* Tenant billing, when your collaborators also have their own S3 storage and have it set as default, will result in their data being stored in their S3 storage.

## Authentication Token

Use the `Create OAuth access token` button to generate an OAuth access token which is valid for **12 hours** after generation. This token can be used by Snowflake and Tableau to access the data in your Base databases and tables for this Project.

See [SnowSQL](https://help.ica.illumina.com/tutorials/base_snowsql#obtaining-oauth-token-and-url) for more information.


# Team

Projects can be shared by updating the project's Team. You can add team members as

* **Existing user** within the current tenant
* By adding their **E-mail** address
* As entire **Workgroup** within the current tenant

Select the corresponding option under **Projects > your\_project > Project Settings > Team > + Add.**

{% hint style="warning" %}
Email invites are sent out as soon as you click the save button on the **add team member** dialog.
{% endhint %}

Users can accept or reject invites. The status column shows a **green checkmark** for accept, an **orange question mark** for users that have not responded and a **red x** for users that rejected the invite.

## Project Owner

The project owner has administrator-level project rights. To change the project owner, select the **Edit project owner** button at the top right and select the new project owner from the list. This can be done by the current project owner, the tenant administrator or a project administrator of the current project.

## Roles

Every user added to the project team will need to have a role assigned for specific categories of functionality in Platform Core. These categories are:

* [Project](/home/h-projects) (contains data and tools to execute analysis)
* [Flow](/project/p-flow) (secondary analysis pipelines)
* [Base](/project/p-base) (genomics data aggregation and analysis)
* [Bench](/project/p-bench) (interactive data analysis)

<figure><img src="/files/eJYinphHqP3FEmjmgz9J" alt="" width="375"><figcaption></figcaption></figure>

{% hint style="info" %}
If a user has been added both as member of a workgroup and as individual user, then the individual rights supersede the group rights. This way, you can add all users in a workgroup and change access for individual users when needed, regardless of their workgroup rights.
{% endhint %}

### Upload and Download rights

While the categories will determine most of what a user can do or see, **explicit upload and download rights need to be granted for users.** Select the checkbox next to Download allowed and Upload allowed when adding a team member.

{% hint style="warning" %}
**Upload and download rights are independent of the assigned role**. A user with only viewer rights will still be able to perform uploads and downloads if their upload and download rights are not disabled. Likewise, an administrator can only perform uploads and downloads if their upload and download rights are enabled.
{% endhint %}

### Project Access

The sections below describe the roles and their allowed actions.

<table><thead><tr><th width="202.30078125"></th><th width="98.01171875" align="center">No Access</th><th width="103.578125" align="center">Data Provider</th><th width="96.36328125" align="center">Viewer</th><th width="123.37890625" align="center">Contributor</th><th width="140.28125" align="center">Administrator</th></tr></thead><tbody><tr><td>Create a Connector</td><td align="center"></td><td align="center">x</td><td align="center">x</td><td align="center">x</td><td align="center">x</td></tr><tr><td>View project resources</td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center">x</td><td align="center">x</td></tr><tr><td>Link/Unlink data to a project</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Subscribe to notifications</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>View Activity</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Create samples</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Delete/archive data</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Manage notification channels</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Manage project team</td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center"></td><td align="center">x</td></tr></tbody></table>

### Flow Access

<table><thead><tr><th width="235.62890625"></th><th width="120.1171875" align="center">No Access</th><th width="99.86328125" align="center">Viewer</th><th width="128.82421875" align="center">Contributor</th></tr></thead><tbody><tr><td>View analyses results</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Create analyses</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Create pipelines and tools</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Edit pipelines and tools</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Add docker image</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr></tbody></table>

### Base Access

<table><thead><tr><th width="195.19140625"></th><th width="125.625" align="center">No Access</th><th width="101.90234375" align="center">Viewer</th><th width="125.83984375" align="center">Contributor</th></tr></thead><tbody><tr><td>View table records</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Click on links in table</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Create queries</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Run queries</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Export query</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Save query</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Export tables</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Create tables</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Load files into a table</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr></tbody></table>

### Bench Access

<table><thead><tr><th></th><th width="124.140625" align="center">No Access</th><th width="123.0703125" align="center">Contributor</th><th width="141.91796875" align="center">Administrator</th></tr></thead><tbody><tr><td>Execute a notebook</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Start/Stop Workspace</td><td align="center"></td><td align="center">x</td><td align="center">x</td></tr><tr><td>Create/Delete/Modify workspaces</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Install additional tools, packages, libraries, …</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Build a new Bench docker image</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>Create a tool for pipeline-execution</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr></tbody></table>


# Connectivity

The platform provides *Connectors* to facilitate automation for operations on data (ie, upload, download, linking).

* [Service connectors](/project/p-connectivity/service-connector) sync data between your local computer or server and the project's cloud-based data storage.
* [Project connectors](/project/p-connectivity/project-connector) link data between individual projects.


# Service Connector

Platform Core provides a Service Connector, which is a small program that runs on your local machine to sync data between the platform's cloud-hosted data store and your local computer or server. The Service Connector securely uploads data or downloads results using TLS 1.2. In order to do this, the Connector makes 2 connections:

* A **control connection**, which the Connector uses to get configuration information from the platform, and to update the platform about its activities
* A **connection towards the storage node**, used to transfer the actual data between your local or network storage and your cloud-based Platform Core storage.

This Connector runs in the background, and configuration is done in the Platform Core UI, where you can add upload and download rules to meet the requirements of the current project and any new projects you may create.

The Service Connector looks at any new files and checks their size. As long as the file size is changing, it knows data is still being added to the file and it is not ready for transfer. Only when the file size is stable and does not change anymore will it consider the file to be complete and initiate transfer. Despite this, it is still best practice to **not connect the Service Connector to active folders which are used as streaming output** for other processes as this can result in incomplete files being transferred when the active processes have extended compute periods in which the file size remains unchanged.

The service connector will handle integrity checking during file transfer, which requires the calculation of hashes on the data. In addition, Transmission speed depends on the available data transfer bandwidth and connection stability. For these reasons, uploading large amounts of data can take considerable time. This can in turn result in temporarily seeing empty folders at the destination location since these are created at the beginning of the transfer process.

{% hint style="info" %}
Both the CLI and the service connector require x86 architecture. For ARM-based architecture on Mac or Windows, you need to keep x86 emulation enabled. Linux does not support x86 emulation.
{% endhint %}

{% tabs %}
{% tab title="Windows" %}
{% embed url="<https://www.youtube.com/embed/PP2OSvEoPug>" %}
{% endtab %}

{% tab title="macOS" %}
{% embed url="<https://www.youtube.com/watch?v=XOGBKT0F-F4>" %}
{% endtab %}

{% tab title="Linux" %}
{% embed url="<https://www.youtube.com/watch?v=YvknZTf7OIs>" %}
{% endtab %}
{% endtabs %}

## Creating a New Connector

1. Select **Projects > your\_project >** **Project Settings > Connectivity > Service Connectors**.
2. Select **+ Create**.
3. Fill out the fields in the New Connector configuration page.
   * Name - Enter the name of the connector.
   * Status - This is automatically updated with the actual status, you do not need to enter anything here.
   * Debug Information Accessible by Illumina (*optional*) - Illumina support can request connector debugging information to help diagnose issues. For security reasons, support can only collect this data if the option *Debug Information Accessible by Illumina* is active. You can choose to either proactively enable this when encountering issues to speed up diagnosis or to only activate it when support requests access. You can at any time revoke access again by deselecting the option.
   * Description (*optional*) - Enter any additional information you want to show for this connector.
   * Mode (*required*) - Specify if the connector can upload data, download data, both or neither.
   * Operating system (*required*) - Select your server or computer operating system.
4. Add any upload or download rules. See [Connector Rules](#connector-rules) below.
5. Select Save and download the connector (top right). An initialization key will be displayed in the platform now. Copy this value as it will be needed during installation.
6. Launch the installer after the download completes and follow the on-screen prompts to complete the installation, including entering the initialization key copied in the previous step. **Do not install the connector in the upload folder** as this will result in the connector attempting to upload itself and the associated log files.

{% tabs %}
{% tab title="Windows" %}

* Run the downloaded .exe file. During the installation, the installer will ask for the initialization key. Fill out the initialization key you see in the platform.
* The installer will create an Illumina Service Connector, register it as a Windows service, and start the service. That means, if you wait for about 60 seconds, and then refresh the screen in the Platform by using the refresh button in the right top corner of the page, the connector should display as connected.
* You can only install 1 connector on Windows. If for some reason, you need to install a new one, first uninstall the old one. You only need to do this when there is a problem with your existing connector. Upgrading a connector is also possible. To do this, you don’t need to uninstall the old one first.
  {% endtab %}

{% tab title="macOS" %}

* Double click the downloaded .dmg file. Double click Illumina Service Connector in the window that opens to start the installer. Run through the installer, and fill out the initialization key when asked for it.
* To start the connector once installed or after a reboot, open the app. You can find the app on the location where you installed it. The connector icon will appear in your dock when the app is running.
* In the platform on the Connectivity page, you can check whether your local connector has been connected with the platform. This can take 60 seconds after you started your connector locally, and you may need to refresh the Connectivity page using the refresh button in the top right corner of the page to see the latest status of your connector.
* The connector app needs to be closed to shut down your computer. You can do this from within your dock.
  {% endtab %}

{% tab title="Linux" %}

* Installations require Java 11 or later. You can check this with ‘java –version’ from a command line terminal. With Java installed, you can run the installer from the command line using the command `bash illumina_unix_develop.sh`.
* Depending on whether you have an X server running or not, it will display a UI, or follow a command line installation procedure. You can force a command line installation by adding a –c flag: `bash illumina_unix_develop.sh -c`.
* The connector can be started by running `./illuminaserviceconnector start` from the folder in which the connector was installed.
  {% endtab %}
  {% endtabs %}

## Connector Rules

In the upload and download rules, you define different properties when setting up a connector. A connector can be used by multiple projects and a connector can have multiple upload and download rules. Configuration can be changed anytime. Changes to the configuration will be applied approximately 60 seconds after changes are made in Platform Core if the connector is already connected. If the connector is not already started when configuration changes are made in Platform Core, it will take about 60 seconds after the connector is started for the configuration changes to be propagated to the connector. The following are the different properties you can configure when setting up a connector. After adding a rule and installing the connector, you can use the **Active** checkbox to disable rules.

Below is an example of a new connector setup with an Upload Rule to find all files ending with `.tar` or `.tar.gz` located within the local folder `C:\Users\username\data\docker-images`.

<figure><img src="/files/KGhOOztAs8X0M5ahUDne" alt=""><figcaption><p>Connector setup</p></figcaption></figure>

### Upload Rules

An upload rule tells the connector which folder on your local disk it needs to watch for new files to upload. The connector contacts the platform every minute to pick up changes to upload rules. To configure upload rules for different projects, first switch into the desired project and select **Connectivity**. Choose the connector from the list and select **Click to add a new upload rule** and define the rule. The project field will be automatically filled with the project you are currently within.

<table><thead><tr><th width="194.17578125">Field</th><th>Description</th></tr></thead><tbody><tr><td>Name</td><td>Name of the upload rule.</td></tr><tr><td>Active</td><td>Set to true to have this rule be active. This allows you to deactivate rules without deleting them.</td></tr><tr><td>Local folder</td><td>The folder path on the local machine where files to be uploaded are stored.</td></tr><tr><td>File pattern</td><td>Files with filenames that match the string/pattern will be uploaded.</td></tr><tr><td>Location</td><td>The location the data will be uploaded to.</td></tr><tr><td>Project</td><td>The project the data will be uploaded to.</td></tr><tr><td>Description</td><td>Additional information about the upload rule.</td></tr><tr><td>Assign Format</td><td>Select which data format tag the uploaded files will receive. This is used for various things like filtering.</td></tr><tr><td>Data owner</td><td>The owner of the data after upload.</td></tr></tbody></table>

### Download Rules

When you schedule downloads in the platform, you can choose which connector needs to download the files. That connector needs some way to know how and where it needs to download your files. That’s what a download rule is for. The connector contacts the platform every minute to pick up changes to download rules. The following are the different download rule settings.

<table><thead><tr><th width="218.046875">Field</th><th>Description</th></tr></thead><tbody><tr><td>Name</td><td>Name of the download rule.</td></tr><tr><td>Active</td><td>Set to true to have this rule be active. This allows you to deactivate rules without deleting them.</td></tr><tr><td>Order of execution</td><td>If using multiple download rules, set the order the rules are performed.</td></tr><tr><td>Target Local folder</td><td>The folder path on the local machine where the files will be downloaded to.</td></tr><tr><td>Description</td><td>Additional information about the download rule.</td></tr><tr><td>Format</td><td>The format the files must comply to in order to be scheduled as downloaded.</td></tr><tr><td>Project</td><td>The projects the rule applies to.</td></tr></tbody></table>

## Connector Status

You can see the service connector status by the color indicator.

<table><thead><tr><th width="125.7265625">Color</th><th width="263.59765625">Status</th></tr></thead><tbody><tr><td>green</td><td>Connected/Active</td></tr><tr><td>orange</td><td>Pending installation</td></tr><tr><td>grey</td><td>Installed/Inactive</td></tr><tr><td>red</td><td>-</td></tr></tbody></table>

## Shared Drives

When you set up your connector for the first time, and your sample files are located on a shared drive, it’s best to create a folder on your local disk, put one of the sample files in there, and do the connector setup with that folder. When this works, try to configure the shared drive.

Transfer to and from a shared drive may be quite slow. That means it can take up to 30 minutes after you configured a shared drive before uploads start. This is due to the integrity check the connector does for each file before it starts uploading. The connector can upload from or download to a shared drive, but there are a few conditions:

* The drive needs to be mounted locally. `X:\illuminaupload` or `/Volumes/shareddrive/illuminaupload` will work, `\\shareddrive\illuminaupload` or `smb://shareddrive/illuminaupload` will not.
* The user running the connector must have access to the shared drive without a password being requested.
* The user who runs the Illumina Service Connector process on the Linux machine needs to have read, write and execute permissions on the installation folder.

## Update connector to newer version

Illumina might release new versions of the Service Connector, with improvements and/or bug fixes. You can easily download a new version of the Connector with the Download button on the Connectivity screen in the platform. After you downloaded the new installer, run it and choose the option ‘Yes, update the existing installation’.

## Uninstall a connector

To uninstall the connector, perform one of the following:

* Windows and Linux: Run the uninstaller located in the folder where the connector was installed.
* Mac: Move the Illumina Service Connector to your Trash folder.

### Log files

The Connector has a log file containing technical information about what’s happening. When something doesn’t work, it often contains clues to why it doesn’t work. Interpreting this log file is not always easy, but it can help the support team to give a fast answer on what is wrong, so it is suggested to attach it to your email when you have upload or download problems. You can find this log file at the following location:

{% tabs %}
{% tab title="Windows" %}
`\<Installation Folder>\logs\BSC.out`\
\
Default: `C:\Program Files (x86)\illumina\logs\BSC.out`
{% endtab %}

{% tab title="macOS" %}
`/<Installation Directory>/Illumina Service Connector.app/Contents/java/app/logs/BSC.out`

\
Default: `/Applications/Illumina Service Connector.app/Contents/java/app/logs/BSC.out`
{% endtab %}

{% tab title="Linux" %}
`/<Installation Directory>/logs/BSC.out`

\
Default: `/usr/local/illumina`
{% endtab %}
{% endtabs %}


# Service Connector Troubleshooting

This topic contains common service connector issues and possible causes.

### General

<table><thead><tr><th width="194.75">General issues</th><th>Solution</th></tr></thead><tbody><tr><td>Connector is connected, but uploads won’t start</td><td><p>To troubleshoot, create a new empty folder on your local disk, add a small file, and configure this folder as upload folder.<br></p><ul><li>If this works, and your sample files are on a <strong>shared drive</strong>, consult the <a href="/pages/BUCLUVW1Yvjj5qp8WxSF#shared-drives">Shared Drives</a> section.</li><li><p>If this works, and your sample files are on a <strong>local disk</strong>, there are several possible causes:</p><ul><li>An error in the platform configuration of the upload folder name.</li><li>For large files, or on slower file systems, the connector needs additional time to start the transfer because it needs to calculate a hash to prevent transfer errors. <strong>Wait up to 30 minutes</strong>, without making changes to your Connector configuration.</li></ul></li><li>If this doesn’t work, you might have a corporate proxy. Proxy configuration is currently not supported for the service connector.</li></ul></td></tr><tr><td>Upload from shared drive does not work</td><td><p>Follow the guidelines in <a href="/pages/BUCLUVW1Yvjj5qp8WxSF#shared-drives">Shared Drives</a> section.<br><br>Inspect the <strong>connector BSC.log</strong> file for any error messages related to the folder not being found.</p><ul><li><p>If there is a related message, possible causes are:</p><ul><li><strong>An issue with the folder name</strong>, such as special characters or spaces. As best practice, use only alphanumeric characters, underscores, dashes and periods.</li><li><strong>A permissions issue</strong>. In this case, ensure the user running the connector has read &#x26; write access, without a password being requested, to the network share.</li></ul></li><li>If there are no messages indicating the folder cannot be found, <strong>wait until the integrity checks are completed</strong>. This check can take considerable time depending on the file systems and network speeds.</li></ul></td></tr><tr><td>Slow Data Transfers</td><td><p>Speed is affected by a number of factors which can change when switching locations (e.g. working from home):</p><ul><li>Distance between upload location and storage location</li><li>Quality of the internet connection (use preferably wired connections)</li><li>Company- or provider-related bandwidth restrictions</li></ul></td></tr><tr><td>Upload or download progress % goes down instead of up.</td><td>This is normal behavior. Instead of one continuous transmission, <strong>data is split into blocks</strong> so if transmission issues occur, not all data has to be retransmit. This can result in dropping back to a lower % of transmission completed when retrying.</td></tr></tbody></table>

### Windows

<table><thead><tr><th width="163.8359375">Issue</th><th>Solution</th></tr></thead><tbody><tr><td>Service connector not connecting</td><td><p>First restart your computer. If that does not solve the problem, open the Services application (enter services.msc in the windows search bar). In there, there should be a <strong>service</strong> called <strong>Illumina Service Connector</strong>.<br></p><ul><li>If the service is not running, start it (right mouse click -> start)</li><li>If the service is running, and still does not connect, you might be on a network with a proxy server. <strong>Proxy configuration is currently not supported</strong> for the connector.</li><li>If you do not have a corporate proxy, and your connector still doesn’t connect, <strong>contact Illumina Technical Support</strong>, and include your connector BSC.out log files.</li></ul></td></tr></tbody></table>

### OSX

<table><thead><tr><th width="163.8359375">Issue</th><th>Solution</th></tr></thead><tbody><tr><td>Service connector not connecting</td><td><p>Check if the Connector is running. If it is, there will be an Illumina icon in your Dock.</p><ul><li>If the service is not running, log out and log back in. An Illumina service connector icon should appear in your dock.</li><li>If not, try starting the Connector manually from the Launchpad menu.</li><li>If the service is running, and still does not connect, you might be on a network with a proxy server. <strong>Proxy configuration is currently not supported</strong> for the connector.</li><li>If you do not have a corporate proxy, and your connector still doesn’t connect, <strong>contact Illumina Technical Support</strong>, and include your connector BSC.out log files.</li></ul></td></tr></tbody></table>

### Linux

<table><thead><tr><th width="163.8359375">Issue</th><th>Solution</th></tr></thead><tbody><tr><td>Corrupted installation script</td><td><p>If you get the following error message <em>“gzip: sfx_archive.tar.gz: not in gzip format. I am sorry, but the installer file seems to be corrupted. If you downloaded that file please try it again. If you transfer that file with ftp please make sure that you are using binary mode.”</em><br></p><p>This indicates the installation script file is corrupted. Please re-download the installation script from Platform Core. Actions like editing the shell script will cause it to be corrupt.</p></td></tr><tr><td>Unsupported version error in log file</td><td>If the log file gives the following error <em>"Unsupported major.minor version 52.0"</em>, an unsupported version of java is present. The connector uses java version 11.</td></tr><tr><td>Manage the connector via the CLI</td><td><p>Connector installation issues:</p><ul><li>Make the connector installation script executable with:<br><code>chmod +x illumina_unix_develop.sh</code></li><li>Once it has been made executable, run the installation script with:<br><code>bash illumina_unix_develop.sh</code></li><li>It may be necessary to run with <code>sudo</code> depending on user permissions on the system:<br><code>sudo bash illumina_unix_develop.sh</code></li><li>If installing on a headless system, use the <code>-c</code> flag to do everything from the command line:<br><code>bash illumina_unix_develop.sh -c</code><br></li><li><strong>Start connector with logging</strong> directly to the terminal stdout) (this should be used when the log file is not present, which can be caused by the absence of java 11). From within the installation directory run:<br><code>./illuminaserviceconnector run</code></li><li><strong>Check status</strong> of connector. From within the install location run:<br><code>./illuminaserviceconnector status</code></li><li><strong>Stop</strong> the connector with:<br><code>./illuminaserviceconnector stop</code></li><li><strong>Restart</strong> the connector with:<br><code>./illuminaserviceconnector restart</code></li></ul></td></tr><tr><td>Can’t define java version for connector</td><td><p>The Service Connector uses Java 11. If you run the installer and encounter the error<br>“Please define INSTALL4J_JAVA_HOME to point to a suitable JVM.”, the installer could not locate a valid Java runtime.</p><p>To resolve the issue, explicitly define the environment variable before running the installation script.</p><p>example:</p><p><code>export INSTALL4J_JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64</code></p><p><code>sh illumina_unix_1_13_2_0_35.sh</code></p><p>Ensure that <code>INSTALL4J_JAVA_HOME</code> points to the <strong>Java installation directory</strong> (the JRE/JDK root), not the java executable.</p><p>Alternatively, <code>INSTALL4J_JAVA_HOME_OVERRIDE</code> can also be used, but using <code>INSTALL4J_JAVA_HOME</code> is recommended</p></td></tr></tbody></table>


# Project Connector

The platform GUI provides the **Project Connector** which allows data to be linked automatically between projects. This creates a **one-way dynamic link** for files and samples from source to destination, meaning that additions and deletions of data in the source project also affect the destination project. This differs from [copying](https://help.ica.illumina.com/project/p-data#copy-data) or [moving](https://help.ica.illumina.com/project/p-data#move-data) which create editable copies of the data.

* **Copy** creates a copy of the file or folder. The data is decoupled and can be edited and deleted.
* **Move** deletes the data on the source destination and moves it to target destination from where it can be edited and deleted. Once deleted, the information ins lost.
* **Manual** **linking** creates a **snapshot** link to files and folder from the source destination. **Changes** to those files at the source are **not propagated** after they have been linked. You can not delete the data, but you can unlink it. to remove it from your destination project.
* **Project Connectors** creates a dynamic link to the source project. All **data which already exists** in the source project before the project connector is created, **is ignored**. All **new data** which is added once the project connector is created, **is automatically propagated to the destination** project. Because this is a dynamic link, **changes to the new data are also propagated**. If there is data propagated which you do not need in the destination project, you can unlink it from there.

<table><thead><tr><th width="114.34765625"></th><th width="84.3203125" align="center">one-way</th><th width="77.90234375" align="center">files</th><th width="97.09765625" align="center">folders</th><th width="124.50390625" align="center">erases source data</th><th width="128.41015625" align="center">propagate source edits</th><th align="center">editable on destination</th></tr></thead><tbody><tr><td>copy</td><td align="center">x</td><td align="center">x</td><td align="center">x</td><td align="center"></td><td align="center"></td><td align="center">x</td></tr><tr><td>move</td><td align="center">x</td><td align="center">x</td><td align="center">x</td><td align="center">x</td><td align="center"></td><td align="center">x</td></tr><tr><td>manual link</td><td align="center">x</td><td align="center">x</td><td align="center">x</td><td align="center"></td><td align="center"></td><td align="center"></td></tr><tr><td>project connector</td><td align="center">x</td><td align="center">x</td><td align="center"></td><td align="center"></td><td align="center">x</td><td align="center"></td></tr></tbody></table>

{% embed url="<https://www.youtube.com/watch?v=tRWdRCWaPU4&ab_channel=Illumina>" %}
Project Connector Setup
{% endembed %}

## Prepare Source Project

1. Select the *source* project (project containing the original data which will be linked to the target project) from the Projects page (**Projects > your\_source\_project**).
2. Select **Project Settings > Details**.
3. Select **Edit**
4. Under **Data Sharing** ensure the value is set to **Yes**
5. Select **Save**

## Creating a New Project Connector

1. Select the *destination* project (the project to which data from the source project will be linked) from the Projects page (**Projects > your\_destination\_project**).
2. From the projects menu, select **Project Settings > Connectivity > Project Connector**
3. Select **+ Create** and complete the necessary fields.
   * Check the box next to **Active** to ensure the connector will be active.
   * **Name** (*required*) — Provide a unique name for the connector.
   * **Type** (*required*) — Select the data type that will be linked (either **File** or **Sample**)
   * **Source Project** - Select the source project where data will be linked from.
   * **Filter Expression** (*optional*) — Enter an expression to restrict which files will be linked via the connector (see [Filter Expression Examples](#filter-expression-examples) below)
   * **Tags** (*optional*) — Add tags to restrict what data will be linked via the connector. Any data in the source project with matching tags will be linked to the destination project.

### Filter Expression Examples

The examples below will restrict linking **Files** based on the **Format** field.

* Only Files with Format of FASTQ will be linked:

  `[?($.details.format.code == 'FASTQ')]`
* Only Files with Format of VCF will be linked:

  `[?($.details.format.code == 'VCF')]`

The examples below will restrict linked **Files** based on a filenames.

* Exact match to 'Sample-1\_S1\_L001\_R1\_001.fastq.gz':

  `[?($.details.name == 'Sample-1_S1_L001_R1_001.fastq.gz')]`
* Ends with '.fastq.gz':

  `[?($.details.name =~ /.*\.fastq.gz/)]`
* Starts with 'Sample-':

  `[?($.details.name =~ /Sample-.*/)]`
* Contains '\_R1\_':

  `[?($.details.name =~ /.*_R1_.*/)]`

The examples below will restrict linking **Samples** based on User Tags and Sample name, respectively.

* Only Samples with the User Tag 'WGS-Project-1'

  `[?('WGS-Project-1' in $.tags.userTags)]`
* Link a Sample with the name 'BSSH\_Sample\_1':

  `[?($.name == 'BSSH_Sample_1')]`


# Notifications

Notifications (**Projects > your\_project > Project Settings > Notifications** ) are events to which you can subscribe. When they are triggered, they deliver a message to an external target system such as emails, Amazon SQS or SNS systems or HTTP post requests. The following table describes available system events to which you can subscribe:

<table><thead><tr><th width="165">Description</th><th width="152">Code</th><th width="266">Details</th><th>Payload</th></tr></thead><tbody><tr><td>Analysis failure</td><td>ICA_EXEC_001</td><td>Emitted when an analysis fails</td><td>Analysis</td></tr><tr><td>Analysis success</td><td>ICA_EXEC_002</td><td>Emitted when an analysis succeeds</td><td>Analysis</td></tr><tr><td>Analysis aborted</td><td>ICA_EXEC_027</td><td>Emitted when an analysis is aborted either by the system or the user</td><td>Analysis</td></tr><tr><td>Analysis status change</td><td>ICA_EXEC_028</td><td>Emitted when an state transition on an analysis occurs</td><td>Analysis</td></tr><tr><td>Base Job failure</td><td>ICA_BASE_001</td><td>Emitted when a Base job fails</td><td>BaseJob</td></tr><tr><td>Base Job success</td><td>ICA_BASE_002</td><td>Emitted when a Base job succeeds</td><td>BaseJob</td></tr><tr><td>Data transfer success</td><td>ICA_DATA_002</td><td>Emitted when a data transfer is marked as Succeeded</td><td>DataTransfer</td></tr><tr><td>Data transfer stalled</td><td>ICA_DATA_025</td><td>Emitted when data transfer hasn't progressed in the past 2 minutes</td><td>DataTransfer</td></tr><tr><td>Data &#x3C;action></td><td>ICA_DATA_100</td><td>Subscribing to this serves as a wildcard for all project data status changes and covers those changes that have no separate code. This does not include DataTransfer events or changes that trigger no data status changes such as adding tags to data.</td><td>ProjectData</td></tr><tr><td>Data linked to project</td><td>ICA_DATA_104</td><td>Emitted when a file is linked to a project</td><td>ProjectData</td></tr><tr><td>Data can not be created in non-indexed folder</td><td>ICA_DATA_105</td><td>Emitted when attempting to create data in a non-indexed folder</td><td>ProjectData</td></tr><tr><td>Data deleted</td><td>ICA_DATA_106</td><td>Emitted when data is deleted</td><td>ProjectData</td></tr><tr><td>Data created</td><td>ICA_DATA_107</td><td>Emitted when data is created</td><td>ProjectData</td></tr><tr><td>Data uploaded</td><td>ICA_DATA_108</td><td>Emitted when data is uploaded</td><td>ProjectData</td></tr><tr><td>Data updated</td><td>ICA_DATA_109</td><td>Emitted when data is updated</td><td>ProjectData</td></tr><tr><td>Data archived</td><td>ICA_DATA_110</td><td>Emitted when data is archived</td><td>ProjectData</td></tr><tr><td>Data unarchived</td><td>ICA_DATA_114</td><td>Emitted when data is unarchived</td><td>ProjectData</td></tr><tr><td>Job status changed</td><td>ICA_JOB_001</td><td>Emitted when a job changes status (INITIALIZED, WAITING_FOR_RESOURCES, RUNNING, STOPPED, SUCCEEDED, PARTIALLY_SUCCEEDED, FAILED)</td><td>JobId</td></tr><tr><td>Sample completed</td><td>ICA_SMP_002</td><td>Emitted when a sample is marked as completed</td><td>ProjectSample</td></tr><tr><td>Sample linked to a project</td><td>ICA_SMP_003</td><td>Emitted when a sample is linked to a project</td><td>ProjectSample</td></tr><tr><td>Workflow session start</td><td>ICA_WFS_001</td><td>Emitted when workflow is started</td><td>WorkflowSession</td></tr><tr><td>Workflow session failure</td><td>ICA_WFS_002</td><td>Emitted when workflow fails</td><td>WorkflowSession</td></tr><tr><td>Workflow session success</td><td>ICA_WFS_003</td><td>Emitted when workflow succeeds</td><td>WorkflowSession</td></tr><tr><td>Workflow session aborted</td><td>ICA_WFS_004</td><td>Emitted when workflow is aborted</td><td>WorkflowSession</td></tr></tbody></table>

When you subscribe to overlapping event codes such as ICA\_EXEC\_002 (analysis success) and ICA\_EXEC\_028 (analysis status change) you will get both notifications when analysis success occurs.

{% hint style="info" %}
When integrating with external systems, it is advised to not solely rely on Platform Core notifications, but to also add a polling system to check the status of long-running tasks. For example verifying the status of long-running (>24h) analyses with a 12 hour interval.
{% endhint %}

## Delivery Targets

Event notifications can be delivered to the following delivery targets:

<table><thead><tr><th width="138">Delivery Target</th><th width="235">Description</th><th>Value</th></tr></thead><tbody><tr><td>Mail</td><td>E-mail delivery</td><td>E-mail Address</td></tr><tr><td>Sqs</td><td>AWS SQS Queue</td><td>AWS SQS Queue URL</td></tr><tr><td>Sns</td><td>AWS SNS Topic</td><td>AWS SNS Topic ARN</td></tr><tr><td>Http</td><td>Webhook (POST request)</td><td>URL</td></tr></tbody></table>

## Subscribing to Notifications

To **create a subscription** via the GUI, select **Projects > your\_project > Project Settings > Notifications** > **+Create > Platform Core event.** Select an event from the dropdown menu and fill out the requested fields. Depending on the selected delivery targets, the fields will change.

<figure><img src="/files/3WScefZuC1dwYDKT0Sm2" alt="" width="563"><figcaption></figcaption></figure>

Once created, you can **disable**, **enable** or **delete** the notification subscriptions at **Projects > your\_project > Project Settings > Notifications**.

{% hint style="info" %}
Subscriptions can only be deleted if there are no failed or pending notifications, so if the delete button is not available, look at the failed notifications details of the subscription. **Projects > your\_project > Project Settings > Notifications > your\_notification > Delivery failed** tab. From there, either reprocess or delete the failed notification so you can delete the notification subscription.
{% endhint %}

### Amazon Resource Policy Settings

In order to allow the platform to deliver events to Amazon SQS or SNS delivery targets, a cross-account policy needs to be added to the target Amazon service.

```json
{
   "Version":"2012-10-17",
   "Statement":[
      {
         "Effect":"Allow",
         "Principal":{
            "AWS":"arn:aws:iam::<platform_aws_account>:root"
         },
         "Action":"<action>",
         "Resource": "<arn>"
      }
   ]
}
```

Substitute the variables in the example above according to the table below.

<table><thead><tr><th width="229.9296875">Variable</th><th>Description</th></tr></thead><tbody><tr><td>platform_aws_account</td><td>The platform AWS account ID: <code>079623148045</code></td></tr><tr><td>action</td><td>For SNS use <code>SNS:Publish</code>. For SQS, use <code>SQS:SendMessage</code></td></tr><tr><td>arn</td><td>The Amazon Resource Name (ARN) of the target SNS topic or SQS queue</td></tr></tbody></table>

See examples for setting policies in [Amazon SQS](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-basic-examples-of-sqs-policies.html) and [Amazon SNS](https://docs.aws.amazon.com/sns/latest/dg/sns-access-policy-use-cases.html).

### Amazon SNS Topic

To create a subscription to deliver events to an Amazon SNS topic, you can use either the GUI or API endpoints.

#### GUI

To create a subscription via the GUI, select **Projects > your\_project > Project Settings > Notifications** > **+Create > Platform Core event.** Select an event from the dropdown menu, insert an optional filter, select the channel type (SNS), and then insert the ARN from the target SNS topic and the AWS region.

<figure><img src="/files/1M2vVt0b65dROXpPZyvw" alt=""><figcaption></figcaption></figure>

#### API

To create a subscription via API, use the endpoint *`/api/notificationChannel`* to create a channel and then `/api/projects/{projectId}/notificationSubscriptions` to create a notification subscription.

### Amazon SQS Queue

To create a subscription to deliver events to an Amazon SQS queue, you can use either GUI or API endpoints.

#### GUI

To create a subscription via the GUI, select **Projects > your\_project > Project Settings > Notifications > +Create > Platform Core event.**

* Select an **event** from the dropdown menu
* Choose **SQS** as the way to receive the notifications and enter your **SQS URL.**
* Depending on the event, you can choose a **payload version**. Not all payload versions are applicable for all events and targets, so the system will filter the options out for you.
* Finally, you can enter a **filter expression** to get only those events are relevant for you. Only those events matching the expression will be received.

#### API

To create a subscription via API, use the endpoint `/api/notificationChannel` to create a channel and then `/api/projects/{projectId}/notificationSubscriptions` to create a notification subscription.

Messages delivered to AWS SQS contain the following event body attributes:

<table><thead><tr><th width="209.91796875">Attribute</th><th>Description</th></tr></thead><tbody><tr><td>correlationId</td><td>GUID used to identify the event</td></tr><tr><td>timestamp</td><td>Date when the event was sent</td></tr><tr><td>eventCode</td><td>Event code of the event</td></tr><tr><td>description</td><td>Description of the event</td></tr><tr><td>payload</td><td>Event payload</td></tr></tbody></table>

The following example is a Data Updated event payload sent to an AWS SQS delivery target (condensed for readability):

```json
{
    "correlationId": "2471d3e2-f3b9-434c-ae83-c7c7d3dcb4e0",
    "timestamp": "2022-10-06T07:51:09.128Z",
    "eventCode": "ICA_DATA_100",
    "description": "Data updates",
    "payload": {
        "id": "fil.8f6f9511d70e4036c60908daa70ea21c",
        ...
    }
}
```

## Filtering

Notification subscriptions will trigger for all events matching the configured event type. A filter may be configured on a subscription to limit the matching strategy to only those event payloads which match the filter.

The filter expressions leverage the [JsonPath](https://github.com/json-path/JsonPath) library for describing the matching pattern to be applied to event payloads. The filter must be in the format `[?(<expression>)]`.

### Examples

The ***Analysis Success*** event delivers a JSON event payload matching the ***Analysis*** data model (as output from the API to [retrieve a project analysis](https://ica.illumina.com/ica/api/swagger/index.html#/Project%20Analysis/getAnalysis)).

```json
 {
     "id": "0c2ed19d-9452-4258-809b-0d676EXAMPLE",
     "timeCreated": "2025-10-16T23:41:04Z",
     "timeModified": "2025-10-17T00:08:00Z",
     "owner":
         {
         "id": "15d51d71-b8a1-4b38-9e3d-74cdfEXAMPLE"
         },
     "tenant":
         {
         "id": "022c9367-8fde-48fe-b129-741a4EXAMPLE",
         "name": "ExampleTenant"
         },
     "reference": "210920-1-CopyToolDev-9d78096d-35f4-47c9-b9b6-e0cbcEXAMPLE",
     "userReference": "210920-1",
     "pipeline":
        {
            "id": "20261676-59ac-4ea0-97bd-8a684EXAMPLE",
            "urn": "urn:ilmn:ica:pipeline:20261676-59ac-4ea0-97bd-8a684EXAMPLE",
            "timeCreated": "2023-07-12T20:03:23Z",
            "timeModified": "2023-07-12T20:03:32Z",
            "owner":
            {
                "id": "15d51d71-b8a1-4b38-9e3d-74cdfEXAMPLE"
            },
            "tenant":
            {
                "id": "022c9367-8fde-48fe-b129-741a4EXAMPLE",
                "name": "ExampleTenant"
            },
            "code": "CopyToolDev",
            "description": "Copy Tool Demonstration",
            "status": "RELEASED",
            "language": "NEXTFLOW",
            "languageVersion":
            {
                "id": "b1585d18-f88c-4ca0-8d47-34f6c01eb6f3",
                "name": "22.04.3",
                "language": "NEXTFLOW"
            },
            "pipelineTags":
            {
                "technicalTags":
                ["Demo"]
            },
            "analysisStorage":
            {
                "id": "6e1b6c8f-f913-4332-9bd0-7fc13eda0fd0",
                "name": "Small",
                "description": "1.2TB"
            },
            "proprietary": false,
            "inputFormType": "XML",
            "reportConfigs":
            {
                "configs":
                []
            }
        },
        "status": "SUCCEEDED",
        "startDate": "2025-10-16T23:41:22Z",
        "endDate": "2025-10-17T00:07:52Z",
        "analysisStorage":
        {
            "id": "6e1b6c8f-f913-4362-9bd0-7fc13eda0fd0",
            "name": "Small",
            "description": "1.2TB"
        },
        "analysisPriority": "HIGH",
        "tags":
        {
            "technicalTags":
            [],
            "userTags":
            [],
            "referenceTags":
            []
        },
        "application":
        {
            "id": "e395cd36-9b1f-4f51-bec0-d4e940cd0739",
            "name": "ICA"
        }
}
```

The below examples demonstrate various filters operating on the above event payload:

* **Filter on a pipeline**, with a **code** that starts with ‘Copy’. You’ll need a regex expression for this:

  `[?($.pipeline.code =~ /Copy.*/)]`
* **Filter on status** (note that the `Analysis success` event is only emitted when the analysis is successful):

  `[?($.status == 'SUCCEEDED')]`

  Both payload Version V3 and V4 guarantee the presence of the final state (SUCCEEDED, FAILED, FAILED\_FINAL, ABORTED) but depending on the flow (so not every intermediate state is guaranteed):

  * V3 can have status REQUESTED - IN\_PROGRESS - SUCCEEDED
  * V4 can have status REQUESTED - QUEUED - INITIALIZING - PREPARING\_INPUTS - IN\_PROGRESS - GENERATING\_OUTPUTS - SUCCEEDED
* **Filter on pipeline**, having a **technical tag** “Demo":

  `[?('Demo' in $.pipeline.pipelineTags.technicalTags)]`
* **Combination of multiple expressions** using `&&`. It's best practice to surround each individual expression with parentheses:

  `[?(($.pipeline.code =~ /Copy.*/) && $.status == 'SUCCEEDED')]`

Examples for other events

* Filtering ICA\_DATA\_104 on owning project name. The top-level keys on which you can filter are under the payload key, so payload is not included in this filter expression.

  `[?($.details.owningProjectName == 'my_project_name')]`

## Custom Events

Custom events let you trigger notification subscriptions using event that are not part of the system-defined event types. When creating a custom subscription, a custom event code can be specified to use within the project. Events can then be sent to the specified event code using a POST API with the request body specifying the event payload.

#### API

Custom events can be defined using the API. To create a custom event for your project, follow the steps below:

1. Create a new custom event `POST {ICA_URL}/ica/rest/api/projects/{projectId}/customEvents`\
   a. Your custom event code must be 1-20 characters long, e.g. 'ICA\_CUSTOM\_123'.\
   b. This event code will be used to reference that custom event type.
2. Create a new notification channel `POST {ICA_URL}/ica/rest/api/notificationChannels`\
   a. If there already is a notification channel with the desired configuration within the same project, you can get the existing channel ID using the call `GET {ICA_URL}/ica/rest/api/notificationChannels`.
3. Create a notification subscription `POST {ICA_URL}/ica/rest/api/projects/{projectId}/customNotificationSubscriptions`.\
   a. Use the event code created in step 1.\
   b. Use the channel ID from step 2.

#### GUI

To create a subscription via the GUI, select **Projects > your\_project > Project Settings > Notifications > +Create > Custom event.**

Once the steps above have been completed successfully, the call from the first step `POST {ICA_URL}/ica/rest/api/projects/{projectId}/customEvents` could be reused with the same event code to continue sending events through the same channel and subscription.

Below is a sample Python function used inside a Platform Core pipeline to post custom events for each failed metric:

```
def post_custom_event(metric_name: str, metric_value: str, threshold: str, sample_name: str):
    api_url = f"{ICA_HOST}/api/projects/{PROJECT_ID}/customEvents"
    headers = {
        "Content-Type": "application/vnd.illumina.v3+json",
        "accept": "application/vnd.illumina.v3+json",
        "X-API-Key": f"{ICA_API_KEY}"
    }
    content = {\"code\": \"ICA_CUSTOM_123\", \"content\": { \"metric_name\": metric_name, \"metric_value\": metric_value,\"threshold\": threshold, \"sample_name\": sample_name}}
    json_data = json.dumps(content)
    response = requests.post(api_url, data=json_data, headers=headers)

    if response.status_code != 204:
        print(f"[EVENT-ERROR] Could not post metric failure event for the metric {metric_name} (sample {sample_name}).")
                
```


# Installation

Download links for the CLI can be found at the [Release History](/command-line-interface/cli-releasehistory).

{% hint style="info" %}
Both the CLI and the service connector require x86 architecture.\
For ARM-based architecture on Mac or Windows, you need to keep x86 emulation enabled.\
Linux does not support x86 emulation.
{% endhint %}

After the file is downloaded, place the CLI in a folder that is included in your $PATH environment variable list of paths, typically /usr/local/bin. Open the Terminal application, navigate to the folder where the downloaded CLI file is located (usually your Downloads folder), and run the following command to copy the CLI file to the appropriate folder. If you do not have write access to your /usr/local/bin folder, then you may use `sudo` (which requires a password) prior to the `cp` command. For example:

{% tabs %}
{% tab title="Mac/Linux" %}
After the file is downloaded, place the CLI in a folder that is included in your $PATH environment variable list of paths, typically /usr/local/bin. Open the Terminal application, navigate to the folder where the downloaded CLI file is located (usually your Downloads folder), and run the following command to copy the CLI file to the appropriate folder. If you do not have write access to your /usr/local/bin folder, then you may use `sudo` (which requires a password) prior to the `cp` command. For example:

```shell
sudo cp icav2 /usr/local/bin
```

If you do not have sudo access on your system, contact your administrator for installation. Alternately, you may place the file in an alternate location and update your $PATH to include this location (see the documentation for your shell to determine how to update this environment variable).

You will also need to make the file executable so that the CLI can run:

```shell
sudo chmod a+x /usr/local/bin/icav2
```

{% endtab %}

{% tab title="Windows" %}
You will likely want to place the CLI in a folder that is included in your $PATH environment variable list of paths. In Windows, you typically want to save your applications in the `C:\Program Files` folder. If you do not have write access to that folder, then open a CMD window in administrative mode (hold down the SHIFT key as you right-click on the CMD application and select "Run as administrator"). Type in the following commands (assuming you have saved ica.exe in your current folder):

```shell
 mkdir "C:\Program Files\Illumina"
 copy icav2.exe "C:\Program Files\Illumina"
```

Then you make sure that the `C:\Program Files\Illumina` folder is included in your `%path%` list of paths. Please do a web search for how to add a path to your %path% system environment variable for your particular version of Windows.
{% endtab %}
{% endtabs %}


# Authentication

The Platform Core CLI uses an Illumina API Key to authenticate. An Illumina API Key can be acquired through the product dashboard after logging into a domain. See [API Keys](/get-started/gs-getstarted#api-keys) for instructions to create an Illumina API Key.

Authenticate using `icav2 config set` command. The CLI will prompt for an `x-api-key` value. Input the API Key generated from the product dashboard here. See the example below (replace `EXAMPLE_API_KEY` with the actual API Key).

```bash
icav2 config set
Creating /Users/johngenome/.icav2/config.yaml
Initialize configuration settings [default]
server-url [ica.illumina.com]: 
x-api-key : EXAMPLE_API_KEY
output-format (allowed values table,yaml,json defaults to table) : 
colormode (allowed values none,dark,light defaults to none) :
```

The CLI will save the API Key to the config file as an encrypted value.

If you want to overwrite existing environment values, use the command `icav2 config set`.\
To remove an existing configuration/session file, use the command `icav2 config reset`.\\

Check the server and confirm you are authenticated using `icav2 config get`

```bash
icav2 config get
access-token: ""
colormode: none
output-format: table
server-url: ica.illumina.com
x-api-key: !!binary HASHED_EXAMPLE_API_KEY
```

If during these steps or in the future you need to reset the authentication, you can do so using the command: `icav2 config reset`


# Data Transfer

The Platform Core CLI can be used for uploading, downloading and viewing information about data stored within Platform Core projects. If not already authenticated, please see the [Authentication](/command-line-interface/cli-authentication) section. Once the CLI has been authenticated with your account, use the command below to list all projects:

`icav2 projects list`

The first column of the output (in default table format) will show the `ID`. This is the **project-id** and will be used in the examples below.

## Upload Data

In this example, we will upload a file called `Sample-1_S1_L001_R1_001.fastq.gz` to the project. Copy your project-id obtained above and use the following command:

`icav2 projectdata upload Sample-1_S1_L001_R1_001.fastq.gz --project-id <project-id>`

To check if the file has uploaded, run the following command to get a list of all files stored within the specified project:

`icav2 projectdata list --project-id <project-id>`

This will show a file ID starting with `fil.` which can be used to get more information about the file and attributes.

`icav2 projectdata get <file-id> --project-id <project-id>`

We have to use `--project-id` in the examples above because we have not entered into a specific project context. To enter a project context use the following command.

`icav2 projects enter <project-name or project-id>`

This will infer the project-id, so that it does not need to be entered for each command.

### Uploading Multiple Files

You can only upload individual files with the `projectdata upload` command. **Wildcards are not supported**.

If you want to upload multiple files in a folder, based on a common name, you can use the following method from the folder where the files are located `ls VAL-0* | xargs -I {} icav2 projectdata upload "{}" /my_upload/my_files/ --project-id <project-id>`

* `ls VAL-0*` lists all files in the current directory whose names start with VAL-O, for example VAL-001.txt, VAL-002.bin,...
* `|` The pipe symbol takes the output of the `ls` command and passes is as input to the next command
* `xargs -I {}` take the list of files and execute the next command while replacing the curly brackets for every individual file.
* `icav2 projectdata upload "{}" /my_upload/my_files/` This command gets executed for each file and uploads that file to the /my\_upload/my\_files/ folder
* `--project-id <project-id>` The id of the project in which you want to upload the files (only needed if you have not entered a project context)

{% hint style="warning" %}
Filenames beginning with / are not allowed, so be careful when entering full path names as those will result in the file being stored on S3 but not being visible in Platform Core. Likewise, folders containing a / in their individual folder name and folders named '.' are not supported
{% endhint %}

## Download Data

The Platform Core CLI can also be used to download files. This can be useful if the download destination is a remote server or HPC cluster into which you are logged in. To download data into the current folder, run the following command:

`icav2 projectdata download <file-id> ./ --project-id <project-id>`

## Temporary Credentials

If the path is provided, the project id from the flag `--project-id` is used. If the `--project-id` flag is not present, then the project id is taken from the context.

The returned AWS credentials for file or folder upload **expire after 12 hours for** [**role-based**](/home/h-storage/s-awss3/iam-role-method) **credentials, 36 hours for all other credentials**.

## Data Transfer Options

For information on options such as using the Platform Core API and AWS CLI to transfer data, visit the [Data Transfer Options](/tutorials/datatransfer) tutorial.


# Config Settings

The Platform Core CLI accepts configuration settings from multiple places, such as environment variables, configuration file, or passed in as command line arguments. When configuration settings are retrieved, the following precedence is used to determine which setting to apply:

1. Command line options - Passed in with the command such as `--access-token`
2. Environment variables - Stored in system environment variables such as `ICAV2_ACCESS_TOKEN`
3. Default config file - Stored by default in the `~/.icav2/config.yaml` on macOS/Linux and `C:\Users\USERNAME\.icav2\.config` on Windows

## Command Line Options

The following global flags are available in the CLI interface:

```bash
-t, --access-token string    JWT used to call rest service
-h, --help                   help for icav2
-o, --output-format string   output format (default "table")
-s, --server-url string      server url to direct commands
-k, --x-api-key string       api key used to call rest service
```

## Environment Variables

Environment variables provide another way to specify configuration settings. Variable names align with the command line options with the following modifications:

* Upper cased
* Prefix `ICAV2_`
* All dashes replaced by underscore

For example, the corresponding environment variable name for the `--access-token` flag is `ICAV2_ACCESS_TOKEN`.

### Disable Retry Mechanism

The environment variable ICAV2\_ICA\_NO\_RETRY\_RATE\_LIMITING allows to disable the retry mechanism. When it is set to `1`, no retries are performed. For any other value, http code 429 will result in 4 retry attempts:

* after 500 milliseconds
* after 2 seconds
* after 10 seconds
* after 30 seconds

## Config File

Upon launching `icav2` for the first time, the configuration yaml file is created and the default config settings are set. Enter an alternative server URL or press enter to leave it as the default. Then enter your API Key and press enter.

After installing the CLI, open a terminal window and enter the `icav2` command. This will initialize a default configuration file in the home folder at `.icav2/config.yaml`.

To reset the configuration, use `./icav2 config reset`

{% hint style="warning" %}
Resetting the configuration removes the configuration from the host device and cannot be undone. The configuration needs to be recreated.
{% endhint %}

Configuration settings are stored in the default configuration file:

```yaml
access-token: ""
colormode: none
output-format: table
server-url: ica.illumina.com
x-api-key: !!binary SMWV6dEXAMPLE
```

The file `~/.icav2/.session.ica.yaml`on macOS/Linux and `C:\Users\USERNAME\.icav2\.session.ica` on Windows will contain the access-token and project-id. These are output files and should not be edited as they are automatically updated.

## Examples

### ICAV2\_X\_API\_KEY

This variable is used to set the [API Key](https://help.ica.illumina.com/get-started/gs-getstarted#api-keys).

1. Command line options - Passed as `--x-api-key <your_api_key>` or `-k <your_api_key>`
2. Environment variables - Stored in system as `ICAV2_X_API_KEY`
3. Default config file - Use `icav2 config set` to update `~/.icav2/config.yaml`(macOS/Linux) or `C:\Users\USERNAME\.icav2\.config` (Windows)


# Output Format

The CLI supports outputs in table, JSON, and YAML formats. The format is set using the `output-format` configuration setting through a command line option, environment variable, or configuration file.

Dates are output as UTC times when using JSON/YAML output format and local times when using table format.

To set the output format, use the following setting:

`--output-format <string>`

* `json` - Outputs in JSON format
* `yaml` - Outputs in YAML format
* `table` - Outputs in tabular format


# Command Index

The build number, together with the used libraries and licenses are provided in the accompanying readme file.

## icav2

```
Command line interface for the Illumina Connected Analytics, a genomics platform-as-a-service

Usage:
  icav2 [command]

Available Commands:
  analysisstorages      Analysis storages commands
  completion            Generate the autocompletion script for the specified shell
  config                Config actions
  dataformats           Data format commands
  help                  Help about any command
  jobs                  Job commands
  metadatamodels        Metadata model commands
  pipelines             Pipeline commands
  projectanalyses       Project analyses commands
  projectdata           Project Data commands
  projectpipelines      Project pipeline commands
  projects              Project commands
  projectsamples        Project samples commands
  regions               Region commands
  storagebundles        Storage bundle commands
  storageconfigurations Storage configurations commands
  tokens                Tokens commands
  version               The version of this application

Flags:
  -t, --access-token string    JWT used to call rest service
  -h, --help                   help for icav2
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -v, --version                version for icav2
  -k, --x-api-key string       api key used to call rest service

Use "icav2 [command] --help" for more information about a command.
```

### icav2 analysisstorages

```
This is the root command for actions that act on analysis storages

Usage:
  icav2 analysisstorages [command]

Available Commands:
  list        list of storage id's

Flags:
  -h, --help   help for analysisstorages

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 analysisstorages [command] --help" for more information about a command.
```

#### icav2 analysisstorages list

```
This command lists all the analysis storage id's

Usage:
  icav2 analysisstorages list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 completion

This command generates custom completion functions for *icav2* tool. These functions facilitate the generation of context-aware suggestions based on the user's input and specific directives provided by the *icav2* tool. For example, for ZSH shell the completion function *\_icav2()* is generated. It could provide suggestions for available commands, flags, and arguments depending on the context, making it easier for the user to interact with the tool without having to constantly refer to documentation.

To enable this custom completion function, you would typically include it in your Zsh configuration (e.g., in .zshrc or a separate completion script) and then use the **compdef** command to associate the function with the icav2 command:

```bash
compdef _icav2 icav2
```

This way, when the user types icav2 followed by a space and presses the TAB key, Zsh will call the \_icav2 function to provide context-aware suggestions based on the user's input and the icav2 tool's directives.

```
Generate the autocompletion script for icav2 for the specified shell.
See each sub-command's help for details on how to use the generated script.

Usage:
  icav2 completion [command]

Available Commands:
  bash        Generate the autocompletion script for bash
  fish        Generate the autocompletion script for fish
  powershell  Generate the autocompletion script for powershell
  zsh         Generate the autocompletion script for zsh

Flags:
  -h, --help   help for completion

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 completion [command] --help" for more information about a command.
```

#### icav2 completion bash

```
Generate the autocompletion script for the bash shell.

This script depends on the 'bash-completion' package.
If it is not installed already, you can install it via your OS's package manager.

To load completions in your current shell session:

	source <(icav2 completion bash)

To load completions for every new session, execute once:

#### Linux:

	icav2 completion bash > /etc/bash_completion.d/icav2

#### macOS:

	icav2 completion bash > $(brew --prefix)/etc/bash_completion.d/icav2

You will need to start a new shell for this setup to take effect.

Usage:
  icav2 completion bash

Flags:
  -h, --help              help for bash
      --no-descriptions   disable completion descriptions

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 completion fish

```
Generate the autocompletion script for the fish shell.

To load completions in your current shell session:

	icav2 completion fish | source

To load completions for every new session, execute once:

	icav2 completion fish > ~/.config/fish/completions/icav2.fish

You will need to start a new shell for this setup to take effect.

Usage:
  icav2 completion fish [flags]

Flags:
  -h, --help              help for fish
      --no-descriptions   disable completion descriptions

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 completion powershell

```
Generate the autocompletion script for powershell.

To load completions in your current shell session:

	icav2 completion powershell | Out-String | Invoke-Expression

To load completions for every new session, add the output of the above command
to your powershell profile.

Usage:
  icav2 completion powershell [flags]

Flags:
  -h, --help              help for powershell
      --no-descriptions   disable completion descriptions

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 completion zsh

```
Generate the autocompletion script for the zsh shell.

If shell completion is not already enabled in your environment you will need
to enable it.  You can execute the following once:

	echo "autoload -U compinit; compinit" >> ~/.zshrc

To load completions in your current shell session:

	source <(icav2 completion zsh)

To load completions for every new session, execute once:

#### Linux:

	icav2 completion zsh > "${fpath[1]}/_icav2"

#### macOS:

	icav2 completion zsh > $(brew --prefix)/share/zsh/site-functions/_icav2

You will need to start a new shell for this setup to take effect.

Usage:
  icav2 completion zsh [flags]

Flags:
  -h, --help              help for zsh
      --no-descriptions   disable completion descriptions

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 config

```
Config command provides functions for CLI configuration management.

Usage:
  icav2 config [command]

Available Commands:
  get         Get configuration information
  reset       Remove the configuration information
  set         Set configuration information

Flags:
  -h, --help   help for config

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 config [command] --help" for more information about a command.
```

#### icav2 config get

```
Get configuration information.

Usage:
  icav2 config get [flags]

Flags:
  -h, --help   help for get
```

#### icav2 config reset

```
Remove configuration information.

Usage:
  icav2 config reset [flags]

Flags:
  -h, --help   help for reset
```

#### icav2 config set

```
Set configuration information. Following information is asked when starting the command : 

 - server-url : used to form the url for the rest api's. 
 - x-api-key : api key used to fetch the JWT used to authenticate to the API server. 
 - colormode : set depending on your background color of your terminal. Input's and errors are colored. Default is 'none', meaning that no colors will be used in the output.
 - table-format : Output layout, defaults to a table, other allowed values are json and yaml

Usage:
  icav2 config set [flags]

Flags:
  -h, --help   help for set
```

### icav2 dataformats

```
This is the root command for actions that act on Data formats

Usage:
  icav2 dataformats [command]

Available Commands:
  list        List data formats

Flags:
  -h, --help   help for dataformats

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 dataformats [command] --help" for more information about a command.
```

#### icav2 dataformats list

```
This command lists the data formats you can use inside of a project

Usage:
  icav2 dataformats list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 help

```
Help provides help for any command in the application.
Simply type icav2 help [path to command] for full details.

Usage:
  icav2 help [command] [flags]

Flags:
  -h, --help   help for help

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 jobs

```
This is the root command for actions that act on jobs

Usage:
  icav2 jobs [command]

Available Commands:
  get         Get details of a job

Flags:
  -h, --help   help for jobs

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 jobs [command] --help" for more information about a command.
```

#### icav2 jobs get

```
This command fetches the details of a job using the argument as an id (uuid).

Usage:
  icav2 jobs get [job id] [flags]

Flags:
  -h, --help   help for get

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 metadatamodels

```
This is the root command for actions that act on metadata models

Usage:
  icav2 metadatamodels [command]

Available Commands:
  list        list of metadata models

Flags:
  -h, --help   help for metadatamodels

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 metadatamodels [command] --help" for more information about a command.
```

#### icav2 metadatamodels list

```
This command lists all the metadata models

Usage:
  icav2 metadatamodels list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 pipelines

```
This is the root command for actions that act on pipelines

Usage:
  icav2 pipelines [command]

Available Commands:
  get         Get details of a pipeline
  list        List pipelines

Flags:
  -h, --help   help for pipelines

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 pipelines [command] --help" for more information about a command.
```

#### icav2 pipelines get

```
This command fetches the details of a pipeline without a project context

Usage:
  icav2 pipelines get [pipeline id] [flags]

Flags:
  -h, --help   help for get

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 pipelines list

```
This command lists the pipelines without the context of a project

Usage:
  icav2 pipelines list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 projectanalyses

```
This is the root command for actions that act on projects analysis

Usage:
  icav2 projectanalyses [command]

Available Commands:
  get         Get the details of an analysis 
  input       Retrieve input of analyses commands
  list        List of analyses for a project 
  output      Retrieve output of analyses commands
  update      Update tags of analyses

Flags:
  -h, --help   help for projectanalyses

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projectanalyses [command] --help" for more information about a command.
```

#### icav2 projectanalyses get

```
This command returns all the details of a analysis.

Usage:
  icav2 projectanalyses get [analysis id] [flags]

Flags:
  -h, --help                help for get
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectanalyses input

```
Retrieve input of analyses commands

Usage:
  icav2 projectanalyses input [analysisId] [flags]

Flags:
  -h, --help                help for input
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectanalyses list

```
This command lists the analyses for a given project. Sorting can be done on 
- reference
- userReference
- pipeline
- status
- startDate
- endDate
- summary

Usage:
  icav2 projectanalyses list [flags]

Flags:
  -h, --help                help for list
      --max-items int       maximum number of items to return, the limit and default is 1000
      --page-offset int     Page offset, only used in combination with sort-by. Offset-based pagination has a result limit of 200K rows and does not guarantee unique results across pages
      --page-size int32     Page size, only used in combination with sort-by. The amount of rows to return. Use in combination with the offset or cursor parameter to get subsequent results. Default and max value of pagesize=1000 (default 1000)
      --project-id string   project ID to set current project context
      --sort-by string      specifies the order to list items

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectanalyses output

```
Retrieve output of analyses commands

Usage:
  icav2 projectanalyses output [analysisId] [flags]

Flags:
  -h, --help                help for output
      --project-id string   project ID to set current project context
      --raw-output          Add this flag if output should be in raw format. Applies only for Cwl pipelines ! This flag needs no value, adding it sets the value to true.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectanalyses update

```
Updates the user and technical tags of an analysis

Usage:
  icav2 projectanalyses update [analysisId] [flags]

Flags:
      --add-tech-tag stringArray      Tech tag to add to analysis. Add flag multiple times for multiple values.
      --add-user-tag stringArray      User tag to add to analysis. Add flag multiple times for multiple values.
  -h, --help                          help for update
      --project-id string             project ID to set current project context
      --remove-tech-tag stringArray   Tech tag to remove from analysis. Add flag multiple times for multiple values.
      --remove-user-tag stringArray   User tag to remove from analysis. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 projectdata

```
This is the root command for actions that act on projects data

Usage:
  icav2 projectdata [command]

Available Commands:
  archive              archive data 
  copy                 Copy data to a project
  create               Create data id for a project
  delete               delete data 
  download             Download a file/folder
  downloadurl          get download url 
  folderuploadsession  Get details of a folder upload
  get                  Get details of a data
  link                 Link data to a project
  list                 List data 
  mount                Mount project data 
  move                 Move data to a project
  temporarycredentials fetch temporal credentials for data
  unarchive            unarchive data 
  unlink               Unlink data to a project
  unmount              Unmount project data 
  update               Updates the details of a data
  upload               Upload a file/folder

Flags:
  -h, --help   help for projectdata

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projectdata [command] --help" for more information about a command.
```

#### icav2 projectdata archive

```
This command archives data for a given project

Usage:
  icav2 projectdata archive [path or data Id] [flags]

Flags:
  -h, --help                help for archive
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata copy

```
This command copies data between projects. Use data id or a combination of path and --source-project-id to identify the source data. By default, the root folder of your current project will be used as destination. If you want to specify a destination, use --destination-folder to specify the destination path or folder id.

Usage:
  icav2 projectdata copy [data id] or [path] [flags]

Flags:
      --action-on-exist string      what to do when a file or folder with the same name already exists: OVERWRITE|SKIP|RENAME (default "SKIP")
      --background                  starts job in background on server. Does not provide upload progress updates. Use icav2 jobs get with the current job.id value
      --copy-instrument-info        copy instrument info form source data to destination data
      --copy-technical-tags         copy technical tags form source data to destination data
      --copy-user-tags              copy user tags form source data to destination data
      --destination-folder string   folder id or path to where you want to copy the data, default root of project
  -h, --help                        help for copy
      --polling-interval int        polling interval in seconds for job status, values lower than 30 will be set to 30 (default 30)
      --project-id string           project ID to set current project context
      --source-project-id string    project ID from where the data needs to be copied, mandatory when using source path notation

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata create

```
This command creates a data on a project. It takes name of file/folder as an argument

Usage:
  icav2 projectdata create [name] [flags]

Flags:
      --data-type string     (*) Data type : FILE or FOLDER
      --folder-id string     Id of the folder
      --folder-path string   Folder path under which the new project data will be created.
      --format string        Only allowed for file, sets the format of the file.
  -h, --help                 help for create
      --project-id string    project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata delete

```
This command deletes data for a given project

Usage:
  icav2 projectdata delete [path or dataId] [flags]

Flags:
  -h, --help                help for delete
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata download

```
Download a file/folder.  Source path can be a data id or a path. Source path for download of a folder should end with '*'. For files : Target defines either local folder into which the download will occur, or a path with a new name for the file. If the  file already exists locally, it is overwritten. For folders : If folder does not exist locally, it will be created automatically. Overwrite of an existing folder will need to be acknowledged.

Usage:
  icav2 projectdata download [source data id or path] [target path] [flags]

Flags:
      --exclude string        Regex filter for file names to exclude from download.
      --exclude-source-path   Indicates that on folder download, the CLI will not create the parent folders of the downloaded folder in ICA on your local machine.
  -h, --help                  help for download
      --project-id string     project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**Example 1**

Using this command all the files starting with **VariantCaller-** will be downloaded (prerequisite: a tool [jq](https://jqlang.github.io/jq/) is installed on the machine):

```bash
icav2 projectdata list --data-type FILE --file-name VariantCaller- --match-mode FUZZY -o json | jq -r '.items[].id' > filelist.txt; for item in $(cat filelist.txt); do echo "--- $item ---"; icav2 projectdata download $item . ; done;
```

**Example 2**

Here an example of how to download all BAM files from a project (we are using some *jq* features to remove '.bam.bai' and '.bam.md5sum' files)

```bash
icav2 projectdata list --file-name .bam --match-mode FUZZY -o json | jq -r '.items[] | select(.details.format.code == "BAM") | [.id] | @tsv' > filelist.txt; for item in $(cat filelist.txt); do echo "--- $item ---"; icav2 projectdata download $item . ; done
```

> Tip: If you want to look up a file id from the GUI, go to that file and open te details view. The file id can be found on the top left side and will begin with fil.

#### icav2 projectdata downloadurl

```
This command returns the data download url for a given project

Usage:
  icav2 projectdata downloadurl [path or data Id] [flags]

Flags:
  -h, --help                help for downloadurl
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata folderuploadsession

```
This command fetches the details a folder upload

Usage:
  icav2 projectdata folderuploadsession [project id] [data id] [folder upload session id] [flags]

Flags:
  -h, --help                help for folderuploadsession
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata get

```
This command fetches the details a data

Usage:
  icav2 projectdata get [data id] or [path] [flags]

Flags:
  -h, --help                help for get
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata link

```
This links data to a project. Use data id or the path + the source project flag identify the data.

Usage:
  icav2 projectdata link [data id] or [path] [flags]

Flags:
  -h, --help                       help for link
      --project-id string          project ID to set current project context
      --source-project-id string   project ID from where the data needs to be linked

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata list

It is best practice to always surround your path with quotes if you want to use the \* wildcard. Otherwise, you may run into situations where the command results in "accepts at most 1 arg(s), received x" as it returns folders with the same name, but different amounts of subfolders.

For more information on how to use pagination, please refer to [Cursor- versus Offset-based Pagination](https://help.ica.illumina.com/reference/r-api#cursor-versus-offset-based-pagination)

If you want to look up a file id from the GUI, go to that file and open te details view. The file id can be found on the top left side and will begin with **fil.**

```
This command lists the data for a given project. Page-offset can only be used in combination with sort-by. Sorting can be done on 
- timeCreated
- timeModified
- name
- path
- fileSizeInBytes
- status
- format
- dataType
- willBeArchivedAt
- willBeDeletedAt

Usage:
  icav2 projectdata list [path] [flags]

Flags:
      --data-type string        Data type. Available values : FILE or FOLDER
      --eligible-link           Add this flag if output should contain only the data that is eligible for linking on the current project. This flag needs no value, adding it sets the value to true.
      --file-name stringArray   The filenames to filter on. The filenameMatchMode-parameter determines how the filtering is done. Add flag multiple times for multiple values.
  -h, --help                    help for list
      --match-mode string       Match mode for the file name. Available values : EXACT (default), EXCLUDE, FUZZY.
      --max-items int           maximum number of items to return, the limit and default is 1000
      --page-offset int         Page offset, only used in combination with sort-by. Offset-based pagination has a result limit of 200K rows and does not guarantee unique results across pages
      --page-size int32         Page size, only used in combination with sort-by. The amount of rows to return. Use in combination with the offset or cursor parameter to get subsequent results. Default and max value of pagesize=1000 (default 1000)
      --parent-folder           Indicates that the given argument is path of the parent folder. All children are selected for list, not the folder itself. This flag needs no value, adding it sets the value to true.
      --project-id string       project ID to set current project context
      --sort-by string          specifies the order to list items
      --status stringArray      Add the status of the data. Available values : PARTIAL, AVAILABLE, ARCHIVING, ARCHIVED, UNARCHIVING, DELETING. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**Example to list files in the folder SOURCE**

```bash
icav2 projectdata list --project-id <project_id> --parent-folder /SOURCE/
```

**Example to list only subfolders in the folder SOURCE**

```bash
icav2 projectdata list --project-id <project_id> --parent-folder /SOURCE/ --data-type FOLDER
```

#### icav2 projectdata mount

```
This command mounts the project data as a file system directory for a given project

Usage:
  icav2 projectdata mount [mount directory path] [flags]

Flags:
      --allow-other         Allow other users to access this project
  -h, --help                help for mount
      --list                List currently mounted projects
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata move

```
This command moves data between projects. Use data id or a combination of path and --source-project-id to identify the source data. By default, the root folder of your current project will be used as destination. If you want to specify a destination, use --destination-folder to specify the destination path or folder id.

Usage:
  icav2 projectdata move [data id] or [path] [flags]

Flags:
      --background                  starts job in background on server. Does not provide upload progress updates. Use icav2 jobs get with the current job.id value
      --destination-folder string   folder id or path to where you want to move the data, default root of project
  -h, --help                        help for move
      --polling-interval int        polling interval in seconds for job status, values lower than 30 will be set to 30 (default 30)
      --project-id string           project ID to set current project context
      --source-project-id string    project ID from where the data needs to be moved, mandatory when using source path notation

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata temporarycredentials

```
This command fetches  temporal AWS and Rclone credentials for a given project-data. If path is given, project id from the flag --project-id is used. If flag not present project is taken from the context

Usage:
  icav2 projectdata temporarycredentials [path or data Id] [flags]

Flags:
  -h, --help                help for temporarycredentials
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata unarchive

```
This command unarchives data for a given project

Usage:
  icav2 projectdata unarchive [path or dataId] [flags]

Flags:
  -h, --help                help for unarchive
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata unlink

```
This unlinks data from a project. Use path or id to identifiy the data.

Usage:
  icav2 projectdata unlink [data id] or [path] [flags]

Flags:
  -h, --help                help for unlink
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata unmount

```
This command unmounts previously mounted project data

Usage:
  icav2 projectdata unmount [flags]

Flags:
      --directory-path string   Set path to unmount
  -h, --help                    help for unmount
      --project-id string       project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata update

```
This command updates some details of a data. Only user/tech tags, format and dates of will be archived/delete can be updated.

Usage:
  icav2 projectdata update [data id] or [path] [flags]

Flags:
      --add-tech-tag stringArray      Tech tag to add. Add flag multiple times for multiple values.
      --add-user-tag stringArray      User tag to add. Add flag multiple times for multiple values.
      --format-code string            Format to assign to the data. Only available for files.
  -h, --help                          help for update
      --project-id string             project ID to set current project context
      --remove-tech-tag stringArray   Tech tag to remove. Add flag multiple times for multiple values.
      --remove-user-tag stringArray   User tag to remove. Add flag multiple times for multiple values.
      --will-be-archived-at string    Time when data will be archived. Format is YYYY-MM-DD. Time is set to 00:00:00UTC time. Only available for files.
      --will-be-deleted-at string     Time when data will be deleted. Format is YYYY-MM-DD. Time is set to 00:00:00UTC time. Only available for files.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectdata upload

```
Upload a file/folder.  For files : if the target path does not already exist, it will be created automatically. For folders : overwrite will need to be acknowledged. Argument "icapath" is optional.

Usage:
  icav2 projectdata upload [local path] [icapath] [flags]

Flags:
      --existing-sample                    Link to existing sample
  -h, --help                               help for upload
      --new-sample                         Create and link to new sample
      --num-workers int                    number of workers to parallelize.  Default calculated based on CPUs available.
      --project-id string                  project ID to set current project context
      --sample-description string          Set Sample Description for new sample
      --sample-id string                   Set Sample id of existing sample
      --sample-name string                 Set Sample name for new sample or from existing sample
      --sample-technical-tag stringArray   Set Sample Technical tag for new sample
      --sample-user-tag stringArray        Set Sample User tag for new sample

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**Example for uploading multiple files**

In this example all the fastq.gz files from *source* will be uploaded to *target* using **xargs** utility.

```bash
find $source -name '*.fastq.gz' | xargs -n 1 -P 10 -I {} icav2 projectdata upload {} /$target/
```

**Example for uploading multiple files using a CSV file**

In this example we upload multiple bam files specified with the corresponding path in the file *bam\_files.csv*. The files will be renamed. We are using screen in detached mode (this creates a new session but not attaching to it):

```bash
while IFS=, read -r current_bam_file_name bam_path new_bam_file_name
do
screen -d -m icav2 projectdata upload ${bam_path}/${current_bam_file} /bam_files/${new_bam_file_name} --project-id $projectID
done <./bam_files.csv 2>./log.txt
```

### icav2 projectpipelines

```
This is the root command for actions that act on projects pipeline

Usage:
  icav2 projectpipelines [command]

Available Commands:
  create      Create a pipeline
  input       Retrieve input parameters of pipeline
  link        Link pipeline to a project
  list        List of pipelines for a project 
  start       Start a pipeline
  unlink      Unlink pipeline from a project

Flags:
  -h, --help   help for projectpipelines

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projectpipelines [command] --help" for more information about a command.
```

#### icav2 projectpipelines create

```
This command creates a  pipeline in the current project

Usage:
  icav2 projectpipelines create [command]

Available Commands:
  cwl          Create a cwl pipeline
  cwljson      Create a cwl Json pipeline
  nextflow     Create a nextflow pipeline
  nextflowjson Create a nextflow Json pipeline

Flags:
  -h, --help                help for create
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projectpipelines create [command] --help" for more information about a command.
```

**icav2 projectpipelines create cwl**

```
This command creates a CWL pipeline in the current project using the argument as code for the pipeline

Usage:
  icav2 projectpipelines create cwl [code] [flags]

Flags:
      --category stringArray   Category of the cwl pipeline. Add flag multiple times for multiple values.
      --comment string         Version comments
      --description string     (*) Description of pipeline
  -h, --help                   help for cwl
      --html-doc string        Html documentation for the cwl pipeline
      --links string           links in json format
      --parameter string       (*) Path to the parameter XML file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --project-id string      project ID to set current project context
      --proprietary            Add the flag if this pipeline is proprietary
      --storage-size string    (*) Name of the storage size. Can be fetched using the command 'icav2 analysisstorages list'.
      --tool stringArray       Path to the tool cwl file. Add flag multiple times for multiple values. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --workflow string        (*) Path to the workflow cwl file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**icav2 projectpipelines create cwljson**

```
This command creates a CWL Json pipeline in the current project using the argument as code for the pipeline

Usage:
  icav2 projectpipelines create cwljson [code] [flags]

Flags:
      --category stringArray         Category of the cwl pipeline. Add flag multiple times for multiple values.
      --comment string               Version comments
      --description string           (*) Description of pipeline
  -h, --help                         help for cwljson
      --html-doc string              Html documentation for the cwl pipeline
      --inputForm string             (*) Path to the input form file.
      --links string                 links in json format
      --onRender string              Path to the on render file.
      --onSubmit string              Path to the on submit file.
      --otherInputForm stringArray   Path to the other input form files. Add flag multiple times for multiple values. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --project-id string            project ID to set current project context
      --proprietary                  Add the flag if this pipeline is proprietary
      --storage-size string          (*) Name of the storage size. Can be fetched using the command 'icav2 analysisstorages list'.
      --tool stringArray             Path to the tool cwl file. Add flag multiple times for multiple values. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --workflow string              (*) Path to the workflow cwl file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**icav2 projectpipelines create nextflow**

```
This command creates a Nextflow pipeline in the current project

Usage:
  icav2 projectpipelines create nextflow [code] [flags]

Flags:
      --category stringArray      Category of the nextflow pipeline. Add flag multiple times for multiple values.
      --comment string            Version comments
      --config string             Path to the config nextflow file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --description string        (*) Description of pipeline
  -h, --help                      help for nextflow
      --html-doc string           Html documentation fo the nexflow pipeline
      --links string              links in json format
      --main string               (*) Path to the main nextflow file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --nextflow-version string   Version of nextflow language to use.
      --other stringArray         Path to the other nextflow file. Add flag multiple times for multiple values. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --parameter string          (*) Path to the parameter XML file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --project-id string         project ID to set current project context
      --proprietary               Add the flag if this pipeline is proprietary
      --storage-size string       (*) Name of the storage size. Can be fetched using the command 'icav2 analysisstorages list'.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**icav2 projectpipelines create nextflowjson**

```
This command creates a Nextflow Json pipeline in the current project

Usage:
  icav2 projectpipelines create nextflowjson [code] [flags]

Flags:
      --category stringArray         Category of the nextflow pipeline. Add flag multiple times for multiple values.
      --comment string               Version comments
      --config string                Path to the config nextflow file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --description string           (*) Description of pipeline
  -h, --help                         help for nextflowjson
      --html-doc string              Html documentation fo the nexflow pipeline
      --inputForm string             (*) Path to the input form file.
      --links string                 links in json format
      --main string                  (*) Path to the main nextflow file. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --nextflow-version string      Version of nextflow language to use.
      --onRender string              Path to the on render file.
      --onSubmit string              Path to the on submit file.
      --other stringArray            Path to the other nextflow file. Add flag multiple times for multiple values. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --otherInputForm stringArray   Path to the other input form files. Add flag multiple times for multiple values. You can set a custom file name and path by adding ':filename=' and the filename with optionally the path the file should be located in.
      --project-id string            project ID to set current project context
      --proprietary                  Add the flag if this pipeline is proprietary
      --storage-size string          (*) Name of the storage size. Can be fetched using the command 'icav2 analysisstorages list'.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectpipelines input

```
Retrieve input parameters of pipeline

Usage:
  icav2 projectpipelines input [pipelineId] [flags]

Flags:
  -h, --help                help for input
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectpipelines link

```
This links a pipeline to a project. Use code or id to identifiy the pipeline. If code is not found, argument is used as id.

Usage:
  icav2 projectpipelines link [pipeline code] or [pipeline id] [flags]

Flags:
  -h, --help                       help for link
      --project-id string          project ID to set current project context
      --source-project-id string   project ID from where the pipeline needs to be linked, mandatory when using pipeline code

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectpipelines list

```
This command lists the pipelines for a given project

Usage:
  icav2 projectpipelines list [flags]

Flags:
  -h, --help                help for list
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectpipelines start

```
This command starts a  pipeline in the current project

Usage:
  icav2 projectpipelines start [command]

Available Commands:
  cwl          Start a CWL pipeline
  cwljson      Start a CWL Json pipeline
  nextflow     Start a Nextflow pipeline
  nextflowjson Start a Nextflow Json pipeline

Flags:
  -h, --help                help for start
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projectpipelines start [command] --help" for more information about a command.
```

**icav2 projectpipelines start cwl**

```
This command starts a CWL pipeline for a given pipeline id, or for a pipeline code from the current project.

Usage:
  icav2 projectpipelines start cwl [pipeline id] or [code] [flags]

Flags:
      --data-id stringArray           Enter data id's as follows : dataId{optional-mount-path} . Add flag multiple times for multiple values.  Mount path is optional and can be absolute and relative and can not contain curly braces.
      --data-parameters stringArray   Enter data-parameters as follows : parameterCode:referenceDataId . Add flag multiple times for multiple values.
  -h, --help                          help for cwl
      --idempotency-key string        Add a maximum 255 character idempotency key to prevent duplicate requests. The  response is retained for 7 days so the key must be unique during that timeframe.
      --input stringArray             Enter inputs as follows : parametercode:dataId,dataId{optional-mount-path},dataId,... . Add flag multiple times for multiple values. Mount path is optional and can be absolute and relative and can not contain curly braces and commas.
      --input-json string             Analysis input JSON string. JSON input works only with file-based CWL pipelines (built using code, not a graphical editor in ICA).
      --output-parent-folder string   The id of the folder in which the output folder should be created.
      --parameters stringArray        Enter single-value parameters as code:value. Enter multi-value parameters as code:"'value1','value2','value3'". To add multiple values, add the flag multiple times.
      --project-id string             project ID to set current project context
      --reference-tag stringArray     Reference tag. Add flag multiple times for multiple values.
      --storage-size string           (*) Name of the storage size. Can be fetched using the command 'icav2 analysisstorages list'
      --technical-tag stringArray     Technical tag. Add flag multiple times for multiple values.
      --type-input string             (*) Input type STRUCTURED or JSON
      --user-reference string         (*) User reference
      --user-tag stringArray          User tag. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**icav2 projectpipelines start cwljson**

```
This command starts a CWL Json pipeline for a given pipeline id, or for a pipeline code from the current project. See ICA CLI documentation for more information (https://help.ica.illumina.com/).

Usage:
  icav2 projectpipelines start cwljson [pipeline id] or [code] [flags]

Flags:
      --field stringArray             Fields. Add flag multiple times for multiple fields. --field fieldA:value --field multivalueFieldB:value1,value2
      --field-data stringArray        Data fields. Add flag multiple times for multiple fields. --field-data fieldA:fil.id --field-data multivalueFieldB:fil.id1,fil.id2
      --group stringArray             Groups. Add flag multiple times for multiple fields in the group. --group groupA.index1.multivalueFieldA:value1,value2 --group groupA.index1.fieldB:value --group groupB.index1.fieldA:value --group groupB.index2.fieldA:value
      --group-data stringArray        Data groups. Add flag multiple times for multiple fields in the group. --group-data groupA.index1.multivalueFieldA:fil.id1,fil.id2 --group-data groupA.index1.fieldB:fil.id --group-data groupB.index1.fieldA:fil.id --group-data groupB.index2.fieldA:fil.id
  -h, --help                          help for cwljson
      --idempotency-key string        Add a maximum 255 character idempotency key to prevent duplicate requests. The  response is retained for 7 days so the key must be unique during that timeframe.
      --output-parent-folder string   The id of the folder in which the output folder should be created.
      --project-id string             project ID to set current project context
      --reference-tag stringArray     Reference tag. Add flag multiple times for multiple values.
      --storage-size string           (*) Name of the storage size. Can be fetched using the command 'icav2  list'.
      --technical-tag stringArray     Technical tag. Add flag multiple times for multiple values.
      --user-reference string         (*) User reference
      --user-tag stringArray          User tag. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**Field definition**

A field can only have values (--field) and a data field can only have datavalues (--field-data). To create multiple fields or data fields, you have to repeat the flag.

For example

```
--field fieldA:valueA --fieldB multivalueFieldB:valueB1,valueB2 --field-data DataFieldC:file.id"
```

matches

```bash
    "fields": [
      {
        "id": "fieldA",
        "values": [
          "valueA"
        ]
      },
      {
        "id": "multivalueFieldB",
        "values": [
          "valueB1",
          "valueB2"
        ]
      },
      {
        "id": "DataFieldC",
        "values": [
          "file.id"
        ]
      }
    ]
```

The following example with --field and --field-data

```
--field asection:SECTION1
--field atext:"this is atext text"
--field ttt:tb1
--field notallowedrole:f
--field notallowedcondition:"this is a not allowed text box"
--field maxagesum:20
--field-data txts1:fil.ade9bd0b6113431a2de108d9fe48a3d8
--field-data txts2:fil.ade9bd0b6113431a2de108d9fe48a3d7{/dir1/dir2},fil.ade9bd0b6113431a2de108d9fe48a3d6{/dir3/dir4}
```

matches

```bash
    "fields": [
    {
      "id": "asection",
      "values": [
        "SECTION1"
      ]
    },
    {
      "id": "atext",
      "values": [
        "this is atext text"
      ]
    },
    {
      "id": "ttt",
      "values": [
        "tb1"
      ]
    },
    {
      "id": "notallowedrole",
      "values": [
        "f"
      ]
    },
    {
      "id": "notallowedcondition",
      "values": [
        "this is a not allowed text box"
      ]
    },
    {
      "id": "maxagesum",
      "values": [
        "20"
      ]
    },
    {
      "dataValues": [
        {
          "dataId": "fil.ade9bd0b6113431a2de108d9fe48a3d8"
        }
      ],
      "id": "txts1"
    },
    {
      "dataValues": [
        {
          "dataId": "fil.ade9bd0b6113431a2de108d9fe48a3d7",
          "mountPath": "/dir1/dir2"
        },
        {
          "dataId": "fil.ade9bd0b6113431a2de108d9fe48a3d6",
          "mountPath": "/dir3/dir4"
        }
      ],
      "id": "txts2"
    }
  ],
```

**Group definition**

A group will only have values (--group) and a data group can only have datavalues (--group-data). Add flags multiple times for multiple groups and fields in the group.

```
--group group1.0.age:80
--group group1.0.role:f
--group group1.0.conditions:cancer,covid
--group-data group1.0.info:fil.a4f17ecf13ca4f692fd008d9fe48a3d7

--group group1.1.age:20
--group group1.1.role:m
--group-data group1.1.info:fil.a4f17ecf13ca4f692fd008d9fe48a3d7

--group group2.0.roleForGroup2:f
```

```bash
"groups": [
    {
      "id": "group1",
      "values": [
        {
          "values": [
            {
              "id": "age",
              "values": [
                "80"
              ]
            },
            {
              "id": "role",
              "values": [
                "f"
              ]
            },
            {
              "id": "conditions",
              "values": [
                "cancer",
                "covid"
              ]
            },
            {
              "dataValues": [
                {
                  "dataId": "fil.a4f17ecf13ca4f692fd008d9fe48a3d7"
                }
              ],
              "id": "info"
            }
          ]
        },
        {
          "values": [
            {
              "id": "age",
              "values": [
                "20"
              ]
            },
            {
              "id": "role",
              "values": [
                "m"
              ]
            },
            {
              "dataValues": [
                {
                  "dataId": "fil.a4f17ecf13ca4f692fd008d9fe48a3d7"
                }
              ],
              "id": "info"
            }
          ]
        }
      ]
    },
    {
      "id": "group2",
      "values": [
        {
          "values": [
            {
              "id": "roleForGroup2",
              "values": [
                "f"
              ]
            }
          ]
        }
      ]
    }
  ]
```

**icav2 projectpipelines start nextflow**

```
This command starts a Nextflow pipeline for a given pipeline id, or for a pipeline code from the current project.

Usage:
  icav2 projectpipelines start nextflow [pipeline id] or [code] [flags]

Flags:
      --data-parameters stringArray   Enter data-parameters as follows : parameterCode:referenceDataId . Add flag multiple times for multiple values.
  -h, --help                          help for nextflow
      --idempotency-key string        Add a maximum 255 character idempotency key to prevent duplicate requests. The  response is retained for 7 days so the key must be unique during that timeframe.
      --input stringArray             Enter inputs as follows : parametercode:dataId,dataId{optional-mount-path},dataId,... . Add flag multiple times for multiple values. Mount path is optional and can be absolute and relative and can not contain curly braces and commas.
      --output-parent-folder string   The id of the folder in which the output folder should be created.
      --parameters stringArray        Enter single-value parameters as code:value. Enter multi-value parameters as code:"'value1','value2','value3'". To add multiple values, add the flag multiple times.
      --project-id string             project ID to set current project context
      --reference-tag stringArray     Reference tag. Add flag multiple times for multiple values.
      --storage-size string           (*) Name of the storage size. Can be fetched using the command 'icav2  list'.
      --technical-tag stringArray     Technical tag. Add flag multiple times for multiple values.
      --user-reference string         (*) User reference
      --user-tag stringArray          User tag. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**icav2 projectpipelines start nextflowjson**

```
This command starts a Nextflow Json pipeline for a given pipeline id, or for a pipeline code from the current project.  See ICA CLI documentation for more information (https://help.ica.illumina.com/).

Usage:
  icav2 projectpipelines start nextflowjson [pipeline id] or [code] [flags]

Flags:
      --field stringArray             Fields. Add flag multiple times for multiple fields. --field fieldA:value --field multivalueFieldB:value1,value2
      --field-data stringArray        Data fields. Add flag multiple times for multiple fields. --field-data fieldA:fil.id --field-data multivalueFieldB:fil.id1,fil.id2
      --group stringArray             Groups. Add flag multiple times for multiple fields in the group. --group groupA.index1.multivalueFieldA:value1,value2 --group groupA.index1.fieldB:value --group groupB.index1.fieldA:value --group groupB.index2.fieldA:value
      --group-data stringArray        Data groups. Add flag multiple times for multiple fields in the group. --group-data groupA.index1.multivalueFieldA:fil.id1,fil.id2 --group-data groupA.index1.fieldB:fil.id --group-data groupB.index1.fieldA:fil.id --group-data groupB.index2.fieldA:fil.id
  -h, --help                          help for nextflowjson
      --idempotency-key string        Add a maximum 255 character idempotency key to prevent duplicate requests. The  response is retained for 7 days so the key must be unique during that timeframe.
      --output-parent-folder string   The id of the folder in which the output folder should be created.
      --project-id string             project ID to set current project context
      --reference-tag stringArray     Reference tag. Add flag multiple times for multiple values.
      --storage-size string           (*) Name of the storage size. Can be fetched using the command 'icav2  list'.
      --technical-tag stringArray     Technical tag. Add flag multiple times for multiple values.
      --user-reference string         (*) User reference
      --user-tag stringArray          User tag. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

**Field definition**

A field can only have values (--field) and a data field can only have datavalues (--field-data). To create multiple fields or data fields, you have to repeat the flag.

For example

```
--field fieldA:valueA --fieldB multivalueFieldB:valueB1,valueB2 --field-data DataFieldC:file.id"
```

matches

```bash
    "fields": [
      {
        "id": "fieldA",
        "values": [
          "valueA"
        ]
      },
      {
        "id": "multivalueFieldB",
        "values": [
          "valueB1",
          "valueB2"
        ]
      },
      {
        "id": "DataFieldC",
        "values": [
          "file.id"
        ]
      }
    ]
```

The following example with --field and --field-data

```
--field asection:SECTION1
--field atext:"this is atext text"
--field ttt:tb1
--field notallowedrole:f
--field notallowedcondition:"this is a not allowed text box"
--field maxagesum:20
--field-data txts1:fil.ade9bd0b6113431a2de108d9fe48a3d8
--field-data txts2:fil.ade9bd0b6113431a2de108d9fe48a3d7{/dir1/dir2},fil.ade9bd0b6113431a2de108d9fe48a3d6{/dir3/dir4}
```

matches

```bash
    "fields": [
    {
      "id": "asection",
      "values": [
        "SECTION1"
      ]
    },
    {
      "id": "atext",
      "values": [
        "this is atext text"
      ]
    },
    {
      "id": "ttt",
      "values": [
        "tb1"
      ]
    },
    {
      "id": "notallowedrole",
      "values": [
        "f"
      ]
    },
    {
      "id": "notallowedcondition",
      "values": [
        "this is a not allowed text box"
      ]
    },
    {
      "id": "maxagesum",
      "values": [
        "20"
      ]
    },
    {
      "dataValues": [
        {
          "dataId": "fil.ade9bd0b6113431a2de108d9fe48a3d8"
        }
      ],
      "id": "txts1"
    },
    {
      "dataValues": [
        {
          "dataId": "fil.ade9bd0b6113431a2de108d9fe48a3d7",
          "mountPath": "/dir1/dir2"
        },
        {
          "dataId": "fil.ade9bd0b6113431a2de108d9fe48a3d6",
          "mountPath": "/dir3/dir4"
        }
      ],
      "id": "txts2"
    }
  ],
```

**Group definition**

A group will only have values (--group) and a data group can only have datavalues (--group-data). Add flags multiple times for multiple groups and fields in the group.

```
--group group1.0.age:80
--group group1.0.role:f
--group group1.0.conditions:cancer,covid
--group-data group1.0.info:fil.a4f17ecf13ca4f692fd008d9fe48a3d7

--group group1.1.age:20
--group group1.1.role:m
--group-data group1.1.info:fil.a4f17ecf13ca4f692fd008d9fe48a3d7

--group group2.0.roleForGroup2:f
```

```bash
"groups": [
    {
      "id": "group1",
      "values": [
        {
          "values": [
            {
              "id": "age",
              "values": [
                "80"
              ]
            },
            {
              "id": "role",
              "values": [
                "f"
              ]
            },
            {
              "id": "conditions",
              "values": [
                "cancer",
                "covid"
              ]
            },
            {
              "dataValues": [
                {
                  "dataId": "fil.a4f17ecf13ca4f692fd008d9fe48a3d7"
                }
              ],
              "id": "info"
            }
          ]
        },
        {
          "values": [
            {
              "id": "age",
              "values": [
                "20"
              ]
            },
            {
              "id": "role",
              "values": [
                "m"
              ]
            },
            {
              "dataValues": [
                {
                  "dataId": "fil.a4f17ecf13ca4f692fd008d9fe48a3d7"
                }
              ],
              "id": "info"
            }
          ]
        }
      ]
    },
    {
      "id": "group2",
      "values": [
        {
          "values": [
            {
              "id": "roleForGroup2",
              "values": [
                "f"
              ]
            }
          ]
        }
      ]
    }
  ]
```

#### icav2 projectpipelines unlink

```
This unlinks a pipeline from a project. Use code or id to identifiy the pipeline. If code is not found, argument is used as id.

Usage:
  icav2 projectpipelines unlink [pipeline code] or [pipeline id] [flags]

Flags:
  -h, --help                help for unlink
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 projects

```
This is the root command for actions that act on projects

Usage:
  icav2 projects [command]

Available Commands:
  create      Create a project
  enter       Enter project context
  exit        Exit project context
  get         Get details of a project
  list        List projects

Flags:
  -h, --help   help for projects

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projects [command] --help" for more information about a command.
```

#### icav2 projects create

```
This command creates a project.

Usage:
  icav2 projects create [projectname] [flags]

Flags:
      --billing-mode string                Billing mode , defaults to PROJECT (default "PROJECT")
      --data-sharing                       Indicates whether the data and samples created in this project can be linked to other Projects. This flag needs no value, adding it sets the value to true.
  -h, --help                               help for create
      --info string                        Info about the project
      --metadata-model string              Id of the metadata model. 
      --owner string                       Owner of the project. Default is the current user
      --region string                      Region of the project. When not specified : takes a default when there is only 1 region, else a choice will be given.
      --short-descr string                 Short pipelineDescription of the project
      --storage-bundle string              Id of the storage bundle. When not specified : takes a default when there is only 1 bundle, else a choice will be given. 
      --storage-config string              An optional storage configuration id to have self managed storage.
      --storage-config-sub-folder string   Required when specifying a storageConfigurationId. The subfolder determines the object prefix of your self managed storage.
      --technical-tag stringArray          Technical tags for this project. Add flag multiple times for multiple values.
      --user-tag stringArray               User tags for this project. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projects enter

```
This command sets the project context for future commands

Usage:
  icav2 projects enter [projectname] or [project id] [flags]

Flags:
  -h, --help   help for enter

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projects exit

```
This command switches the user back to their personal context

Usage:
  icav2 projects exit [flags]

Flags:
  -h, --help                help for exit
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projects get

```
This command fetches the details of the current project. If no project id is given, we take the one from the config file.

Usage:
  icav2 projects get [project id] [flags]

Flags:
  -h, --help   help for get

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projects list

```
This command lists the projects for the current user. Page-offset can only be used in combination with sort-by. Sorting can be done on 
- name
- shortDescription
- information

Usage:
  icav2 projects list [flags]

Flags:
  -h, --help              help for list
      --max-items int     maximum number of items to return, the limit and default is 1000
      --page-offset int   Page offset, only used in combination with sort-by. Offset-based pagination has a result limit of 200K rows and does not guarantee unique results across pages
      --page-size int32   Page size, only used in combination with sort-by. The amount of rows to return. Use in combination with the offset or cursor parameter to get subsequent results. Default and max value of pagesize=1000 (default 1000)
      --sort-by string    specifies the order to list items

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 projectsamples

```
This is the root command for actions that act on projects samples

Usage:
  icav2 projectsamples [command]

Available Commands:
  complete    Set sample to complete
  create      Create a sample for a project
  delete      Delete a sample for a project
  get         Get details of a sample
  link        Link data to a sample for a project
  list        List of samples for a project 
  listdata    List data from given sample
  unlink      Unlink data from a sample for a project
  update      Update a sample for a project

Flags:
  -h, --help   help for projectsamples

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 projectsamples [command] --help" for more information about a command.
```

#### icav2 projectsamples complete

```
The sample status will be set to 'Available' and a sample completed event will be triggered as well.

Usage:
  icav2 projectsamples complete [sampleId] [flags]

Flags:
  -h, --help                help for complete
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples create

```
This command creates a sample for a project. It takes the name of the sample as argument.

Usage:
  icav2 projectsamples create [name] [flags]

Flags:
      --description string          Description 
  -h, --help                        help for create
      --project-id string           project ID to set current project context
      --technical-tag stringArray   Technical tag. Add flag multiple times for multiple values.
      --user-tag stringArray        User tag. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples delete

```
This command deletes a sample from a project. The different flags indicate the way they are deleted. Only 1 flag can be used.

Usage:
  icav2 projectsamples delete [sampleId] [flags]

Flags:
      --deep         Delete the entire sample: sample and linked files will be deleted from your project.
  -h, --help         help for delete
      --mark         Mark a sample as deleted.
      --unlink       Unlinking the sample: sample is deleted and files are unlinked and available again for linking to another sample.
      --with-input   Delete the sample as well as its input data: sample is deleted from your project, the input files and pipeline output folders are still present in the project but will not be available for linking to a new sample.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples get

```
This command fetches the details a sample using the argument as a name, if nothing found, the argument is used as an id (uuid).

Usage:
  icav2 projectsamples get [sample id] or [name] [flags]

Flags:
  -h, --help                help for get
      --project-id string   project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples link

```
This command adds data to a project sample. Argument is the id of the project sample

Usage:
  icav2 projectsamples link [sampleId] [flags]

Flags:
      --data-id stringArray   (*) Data id of the data that needs to be linked to the project sample. Add flag multiple times for multiple values.
  -h, --help                  help for link
      --project-id string     project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples list

```
This command lists the samples for a given project

Usage:
  icav2 projectsamples list [flags]

Flags:
  -h, --help                        help for list
      --include-deleted             Include the deleted samples in the list. Default set to false.
      --project-id string           project ID to set current project context
      --technical-tag stringArray   Technical tags to filter on. Add flag multiple times for multiple values.
      --user-tag stringArray        User tags to filter on. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples listdata

```
This command lists the data for a given sample. It only supports offset based, and default sorting is done on timeCreated. Sorting can be done on timeCreated
- timeModified
- name
- path
- fileSizeInBytes
- status
- format
- dataType
- willBeArchivedAt
- willBeDeletedAt

Usage:
  icav2 projectsamples listdata [sampleId] [path] [flags]

Flags:
      --data-type string        Data type. Available values : FILE or FOLDER
      --file-name stringArray   The filenames to filter on. The filenameMatchMode-parameter determines how the filtering is done. Add flag multiple times for multiple values.
  -h, --help                    help for listdata
      --match-mode string       Match mode for the file name. Available values : EXACT (default), EXCLUDE, FUZZY.
      --max-items int           maximum number of items to return, the limit and default is 1000
      --page-offset int         Page offset, only used in combination with sort-by. Offset-based pagination has a result limit of 200K rows and does not guarantee unique results across pages
      --page-size int32         Page size, only used in combination with sort-by. The amount of rows to return. Use in combination with the offset or cursor parameter to get subsequent results. Default and max value of pagesize=1000 (default 1000)
      --parent-folder           Indicates that the given argument is path of the parent folder. All children are selected for list, not the folder itself. This flag needs no value, adding it sets the value to true.
      --project-id string       project ID to set current project context
      --sort-by string          specifies the order to list items (default "timeCreated Desc")
      --status stringArray      Add the status of the data. Available values : PARTIAL, AVAILABLE, ARCHIVING, ARCHIVED, UNARCHIVING, DELETING. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples unlink

```
This command removes data from a project sample. Argument is the id of the project sample

Usage:
  icav2 projectsamples unlink [sampleId] [flags]

Flags:
      --data-id stringArray   (*) Data id of the data that will be removed from the project sample. Add flag multiple times for multiple values.
  -h, --help                  help for unlink
      --project-id string     project ID to set current project context

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 projectsamples update

```
This command updates a sample for a project. Name,description, user and technical tags can be updated

Usage:
  icav2 projectsamples update [sampleId] [flags]

Flags:
      --add-tech-tag stringArray      Tech tag to add. Add flag multiple times for multiple values.
      --add-user-tag stringArray      User tag to add. Add flag multiple times for multiple values.
  -h, --help                          help for update
      --name string                   Name 
      --project-id string             project ID to set current project context
      --remove-tech-tag stringArray   Tech tag to remove. Add flag multiple times for multiple values.
      --remove-user-tag stringArray   User tag to remove. Add flag multiple times for multiple values.

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 regions

```
This is the root command for actions that act on regions

Usage:
  icav2 regions [command]

Available Commands:
  list        list of regions

Flags:
  -h, --help   help for regions

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 regions [command] --help" for more information about a command.
```

#### icav2 regions list

```
This command lists all the regions

Usage:
  icav2 regions list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 storagebundles

```
This is the root command for actions that act on storage bundles

Usage:
  icav2 storagebundles [command]

Available Commands:
  list        list of storage bundles

Flags:
  -h, --help   help for storagebundles

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 storagebundles [command] --help" for more information about a command.
```

#### icav2 storagebundles list

```
This command lists all the storage bundles id's

Usage:
  icav2 storagebundles list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 storageconfigurations

```
This is the root command for actions that act on storage configurations

Usage:
  icav2 storageconfigurations [command]

Available Commands:
  list        list of storage configurations

Flags:
  -h, --help   help for storageconfigurations

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 storageconfigurations [command] --help" for more information about a command.
```

#### icav2 storageconfigurations list

```
This command lists all the storage configurations

Usage:
  icav2 storageconfigurations list [flags]

Flags:
  -h, --help   help for list

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 tokens

```
This is the root command for actions that act on tokens

Usage:
  icav2 tokens [command]

Available Commands:
  create      Create a JWT token
  refresh     Refresh a JWT token from basic authentication

Flags:
  -h, --help   help for tokens

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service

Use "icav2 tokens [command] --help" for more information about a command.
```

#### icav2 tokens create

```
This command creates a JWT token from the API key.

Usage:
  icav2 tokens create [flags]

Flags:
  -h, --help   help for create

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

#### icav2 tokens refresh

```
This command refreshes a JWT token from basic authentication with gantype JWT-bearer that is set with the -t flag.

Usage:
  icav2 tokens refresh [flags]

Flags:
  -h, --help   help for refresh

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```

### icav2 version

```
The version of this application

Usage:
  icav2 version [flags]

Flags:
  -h, --help   help for version

Global Flags:
  -t, --access-token string    JWT used to call rest service
  -o, --output-format string   output format (default "table")
  -s, --server-url string      server url to direct commands
  -k, --x-api-key string       api key used to call rest service
```


# Releases

Find the links to CLI builds in the Releases section below.

## Downloading the Installer

In the [Releases section](#releases) below, select the matching **operating system** in the **link column** for the version you want to install. This will download the installer for that operating system.

## Version Check

To determine which **CLI version** you are **currently using**, navigate to your currently installed CLI and use the CLI command `icav2 version` For help on this command use `icav2 version -h`.

## Integrity Check

Checksums are provided alongside each downloadable CLI binary to verify file integrity. The checksums are generated using the SHA256 algorithm. To use the checksums:

1. Download the CLI binary for your OS
2. Download the corresponding checksum using the links in the table
3. Calculate the SHA256 checksum of the downloaded CLI binary
4. Diff the calculated SHA256 checksum with the downloaded checksum. If the checksums match, the integrity is confirmed.

There are a variety of open source tools for calculating the SHA256 checksum. See the below tables for examples.

For CLI v2.3.0 and later:

| OS      | Command                                           |
| ------- | ------------------------------------------------- |
| Windows | `CertUtil -hashfile ica-windows-amd64.zip SHA256` |
| Linux   | `sha256sum ica-linux-amd64.zip`                   |
| Mac     | `shasum -a 256 ica-darwin-amd64.zip`              |

<details>

<summary>For CLI v2.2.0</summary>

| OS      | Command                               |
| ------- | ------------------------------------- |
| Windows | `CertUtil -hashfile icav2.exe SHA256` |
| Linux   | `sha256sum icav2`                     |
| Mac     | `shasum -a 256 icav2`                 |

</details>

## Releases

| Version | Link                                                                                                        | Checksum                                                                                                      |
| ------- | ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| 2.47    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.47.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.47.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.47.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.47.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.47.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.47.0/ica-linux-amd64.sha256)   |
| 2.46    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.46.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.46.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.46.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.46.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.46.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.46.0/ica-linux-amd64.sha256)   |
| 2.45    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.45.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.45.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.45.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.45.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.45.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.45.0/ica-linux-amd64.sha256)   |
| 2.44    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.44.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.44.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.44.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.44.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.44.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.44.0/ica-linux-amd64.sha256)   |
| 2.43    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.43.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.43.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.43.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.43.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.43.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.43.0/ica-linux-amd64.sha256)   |
| 2.42    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.42.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.42.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.42.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.42.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.42.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.42.0/ica-linux-amd64.sha256)   |
| 2.41    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.41.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.41.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.41.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.41.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.41.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.41.0/ica-linux-amd64.sha256)   |
| 2.40    | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.40.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.40.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.40.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.40.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.40.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.40.0/ica-linux-amd64.sha256)   |
| 2.39.0  | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.39.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.39.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.39.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.39.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.39.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.39.0/ica-linux-amd64.sha256)   |
| 2.38.0  | No changes, use 2.37.0                                                                                      | -                                                                                                             |
| 2.37.0  | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.37.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.37.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.37.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.37.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.37.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.37.0/ica-linux-amd64.sha256)   |
| 2.36.0  | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.36.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.36.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.36.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.36.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.36.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.36.0/ica-linux-amd64.sha256)   |
| 2.35.0  | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.35.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.35.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.35.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.35.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.35.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.35.0/ica-linux-amd64.sha256)   |
| 2.34.0  | [Mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.34.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.34.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.34.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.34.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.34.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.34.0/ica-linux-amd64.sha256)   |
| 2.33.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.33.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.33.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.33.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.33.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.33.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.33.0/ica-linux-amd64.sha256)   |
| 2.32.2  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.32.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.32.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.32.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.32.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.32.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.32.0/ica-linux-amd64.sha256)   |
| 2.31.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.31.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.31.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.31.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.31.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.31.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.31.0/ica-linux-amd64.sha256)   |
| 2.30.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.30.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.30.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.30.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.30.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.30.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.30.0/ica-linux-amd64.sha256)   |
| 2.29.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.29.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.29.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.29.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.29.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.29.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.29.0/ica-linux-amd64.sha256)   |
| 2.28.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.28.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.28.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.28.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.28.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.28.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.28.0/ica-linux-amd64.sha256)   |
| 2.27.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.27.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.27.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.27.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.27.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.27.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.27.0/ica-linux-amd64.sha256)   |
| 2.26.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.26.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.26.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.26.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.26.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.26.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.26.0/ica-linux-amd64.sha256)   |
| 2.25.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.25.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.25.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.25.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.25.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.25.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.25.0/ica-linux-amd64.sha256)   |
| 2.24.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.24.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.24.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.24.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.24.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.24.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.24.0/ica-linux-amd64.sha256)   |
| 2.23.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.23.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.23.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.23.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.23.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.23.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.23.0/ica-linux-amd64.sha256)   |
| 2.22.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.22.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.22.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.22.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.22.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.22.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.22.0/ica-linux-amd64.sha256)   |
| 2.21.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.21.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.21.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.21.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.21.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.21.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.21.0/ica-linux-amd64.sha256)   |
| 2.19.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.19.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.19.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.19.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.19.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.19.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.19.0/ica-linux-amd64.sha256)   |
| 2.18.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.18.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.18.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.18.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.18.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.18.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.18.0/ica-linux-amd64.sha256)   |
| 2.17.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.17.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.17.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.17.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.17.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.17.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.17.0/ica-linux-amd64.sha256)   |
| 2.16.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.16.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.16.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.16.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.16.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.16.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.16.0/ica-linux-amd64.sha256)   |
| 2.15.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.15.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.15.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.15.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.15.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.15.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.15.0/ica-linux-amd64.sha256)   |
| 2.12.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.12.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.12.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.12.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.12.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.12.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.12.0/ica-linux-amd64.sha256)   |
| 2.10.0  | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.10.0/ica-darwin-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.10.0/ica-darwin-amd64.sha256)  |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.10.0/ica-windows-amd64.zip) | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.10.0/ica-windows-amd64.sha256) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.10.0/ica-linux-amd64.zip)     | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.10.0/ica-linux-amd64.sha256)   |
| 2.9.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.9.0/ica-darwin-amd64.zip)       | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.9.0/ica-darwin-amd64.sha256)   |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.9.0/ica-windows-amd64.zip)  | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.9.0/ica-windows-amd64.sha256)  |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.9.0/ica-linux-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.9.0/ica-linux-amd64.sha256)    |
| 2.8.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.8.0/ica-darwin-amd64.zip)       | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.8.0/ica-darwin-amd64.sha256)   |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.8.0/ica-windows-amd64.zip)  | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.8.0/ica-windows-amd64.sha256)  |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.8.0/ica-linux-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.8.0/ica-linux-amd64.sha256)    |
| 2.4.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.4.0/ica-darwin-amd64.zip)       | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.4.0/ica-darwin-amd64.sha256)   |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.4.0/ica-windows-amd64.zip)  | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.4.0/ica-windows-amd64.sha256)  |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.4.0/ica-linux-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.4.0/ica-linux-amd64.sha256)    |
| 2.3.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.3.0/ica-darwin-amd64.zip)       | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.3.0/ica-darwin-amd64.sha256)   |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.3.0/ica-windows-amd64.zip)  | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.3.0/ica-windows-amd64.sha256)  |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.3.0/ica-linux-amd64.zip)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.3.0/ica-linux-amd64.sha256)    |
| 2.2.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.2.0/mac/icav2)                  | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.2.0/mac/icav2_mac.sha)         |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.2.0/windows/icav2.exe)      | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.2.0/windows/icav2_windows.sha) |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.2.0/linux/icav2)              | [sha256](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.2.0/linux/icav2_linux.sha)     |
| 2.1.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.1.0/mac/icav2)                  |                                                                                                               |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.1.0/windows/icav2.exe)      |                                                                                                               |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.1.0/linux/icav2)              |                                                                                                               |
| 2.0.0   | [mac](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.0.0/mac/icav2)                  |                                                                                                               |
|         | [windows](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.0.0/windows/icav2.exe)      |                                                                                                               |
|         | [linux](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/cli/2.0.0/linux/icav2)              |                                                                                                               |

{% hint style="info" %}
To access release history of CLI versions prior to v2.0.0, please see the Illumina Connected Analytics v1 documentation [here](https://illumina.gitbook.io/ica-v1/command-line-interface/cli-releasehistory).
{% endhint %}


# Cloud Analysis Auto-launch

Please see the [Illumina BioInsight Platform](https://help.connected.illumina.com/) site for all content related to Cloud Analysis Auto-Launch:

* [Cloud Analysis Auto-Launch Guidance](https://help.connected.illumina.com/analysis/analysis_autolaunch)
* [Sequencer Auto-launch Analyses Compatibility](https://help.connected.illumina.com/analysis/sequencer-reference)
* [Sample Sheet v2 Guidance](https://help.connected.illumina.com/run-set-up/overview)
* [NovaSeqX: BCL Convert Auto-launch Analysis in Cloud](https://help.connected.illumina.com/cross-product-tutorials/novaseqx-bcl-autolaunch)[ Guided Example](https://help.connected.illumina.com/cross-product-tutorials/novaseqx-bcl-autolaunch)
* [NovaSeq 6000: BCL Convert Auto-launch Analysis in Cloud Guided Example](https://help.connected.illumina.com/cross-product-tutorials/novaseq6000-bcl-autolaunch)


# Nextflow Pipeline

In this tutorial, we will show how to create and launch a pipeline using the Nextflow language in Platform Core.

This tutorial references the [Basic pipeline](https://www.nextflow.io/example1.html) example in the Nextflow documentation.

## Create the pipeline

The first step in creating a pipeline is to create a [Project](https://help.ica.illumina.com/home/h-projects). In the example below, the project is named *Getting Started*.

### Pipeline

After creating your project,

1. **Open the project** at **Projects > your\_project**.
2. Navigate to the **Flow > Pipelines** view in the left navigation pane.
3. From the Pipelines view, click **+Create > Nextflow > XML based** to start creating the Nextflow pipeline.

<figure><img src="/files/SnBYlmQxaVLg7vHt2Cgd" alt=""><figcaption></figcaption></figure>

In the Nextflow pipeline creation view, the Description field is used to add information about the pipeline. Add values for the required Code (unique pipeline name), description and size fields.

<figure><img src="/files/Nq9RgSAKnaPNpK6YYj9C" alt=""><figcaption></figcaption></figure>

### Nextflow Files

Next a Nextflow pipeline definition must be created. The pipeline in this example is a modified version of the Basic pipeline example from the Nextflow documentation.

The description of the pipeline from the linked Nextflow docs:

{% hint style="info" %}
This example shows a pipeline that is made of two processes. The first process receives a FASTA formatted file and splits it into file chunks whose names start with the prefix seq\_.

The process that follows, receives these files and reverses their content by using the rev command line tool.
{% endhint %}

Some modifications are made to the Nextflow pipeline, you do not need to copy these modification by hand. Copyable code is provided below.

* Adding the `container` directive to each process with the desired ubuntu image. If no Docker image is specified, public.ecr.aws/lts/ubuntu:22.04\_stable is used as default. If you want to use the latest image, use *`container 'public.ecr.aws/lts/ubuntu:latest'`*
* Adding the `publishDir` directive with value `'out'` to the `reverse` process.
* Modifying the `reverse` process to write the output to a file `test.txt` instead of stdout.
* Creating a channel with the input file.

**Setting Process Resources**: For each process, you can use the [memory directive](https://www.nextflow.io/docs/latest/process.html#memory) and [cpus directive](https://www.nextflow.io/docs/latest/process.html#cpus) to set the [Compute Types](https://github.com/illumina-swi/ica-docs/blob/stage/docs/tutorials/f-pipelines.md#compute-types). Platform Core will then determine the best matching compute type based on those settings. Suppose you set `memory '10240 GB'` and `cpus 6`, then Platform Core will determine you need the `standard-large` compute type.

Syntax example:

```nf
process iwantstandardsmallresources {
    cpus 2
    memory '8 GB'
    ...
```

Navigate to the **Nextflow files > main.nf** tab to add the definition to the pipeline. Since this is a single file pipeline, we don't need to add any additional definition files. Paste the following definition into the text editor:

```nf
#!/usr/bin/env nextflow
params.in = "$HOME/sample.fa"

// -----------------------------
// Processes
// -----------------------------

// Split the file
process splitSequences {
    container 'public.ecr.aws/lts/ubuntu:latest'

    input:
    path 'input_fa'

    output:
    path "seq_*"

    script:
    """
   awk '/^>/{f="seq_"++d} {print > f}' < ${input_fa}
    """
}

// Reverse the Sequence
process reverse {
    container 'public.ecr.aws/lts/ubuntu:latest'
    publishDir 'out'

    input:
    path x

    output:
    path "test.txt"

    script:
    """
    cat ${x} | rev > test.txt
    """
}

// -----------------------------
// Workflow block
// -----------------------------

workflow {
//     Create a channel with your input file
    sequences = Channel.fromPath(params.in)
    splitSequences(sequences) | reverse | view
}
```

<figure><img src="/files/OvjbQa2uMReHwmgOay0l" alt=""><figcaption></figcaption></figure>

### Input Form

Next create the input form used when launching the pipeline. This is done in the **XML Configuration** tab. Since the pipeline takes in a single FASTA file as input, the input form includes a single file input.

Paste the below XML input form into the XML CONFIGURATION text editor

```xml
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition">
    <pd:dataInputs>
        <pd:dataInput code="in" format="FASTA" type="FILE" required="true" multiValue="false">
            <pd:label>in</pd:label>
            <pd:description>fasta file input</pd:description>
        </pd:dataInput>
    </pd:dataInputs>
    <pd:steps/>
</pd:pipeline>
```

On the left, you see the XML code, on the right, you can see the input form simulation which appears when you use the **simulate** button at the bottom.

<figure><img src="/files/U7tPiuzf74RM9WiCh6Vi" alt=""><figcaption></figcaption></figure>

Once the definition has been added and the input form has been defined, the pipeline is complete.

{% hint style="info" %}
On the **Documentation tab**, you can add additional information about your pipeline. This information will be presented under the Documentation tab whenever a user starts a new analysis on the pipeline.
{% endhint %}

Click the **Save** button at the top right. The pipeline will now be visible from the **Projects > your\_project > Pipelines** view within the project.

<figure><img src="/files/enaEJyikT2k3ZC0dKapx" alt=""><figcaption></figcaption></figure>

## Launch the pipeline

Before launching the pipeline, upload a FASTA file to use as input. For this tutorial, use a public FASTA file from the [UCSC Genome Browser](https://genome.ucsc.edu/). Download [chr1\_GL383518v1\_alt.fa.gz](https://hgdownload.cse.ucsc.edu/goldenpath/hg38/chromosomes/chr1_GL383518v1_alt.fa.gz) and unzip yjr FASTA file to decompress it.

To upload the FASTA file to the project, navigate to **Projects > your\_project > Data**. In the Data view, drag and drop the FASTA file from your local machine in the input section (2) in the browser. Once the file upload completes, the file record will show in the Data explorer. The file format should be auto-detected and be FASTA. If this is not the case, you can set it by hand by selecting the file and changing the format from the manage menu item.

<figure><img src="/files/8T2EHzLlDCATqHrqpnWX" alt=""><figcaption></figcaption></figure>

Now that the input data is uploaded, we can proceed to launch the pipeline. Navigate to **Projects > your\_project > Flow > Analyses** click on **Start**. Next, select your pipeline from the list.

{% hint style="info" %}
Alternatively you can start your pipeline from **Projects > your\_project > Flow > Pipelines > your\_pipeline > Start analysis**.
{% endhint %}

In the Launch Pipeline view, the input form fields are shown along with some required information to create the analysis.

With the required information set, click **Start Analysis**.

## Monitoring Analysis

After launching the pipeline, navigate to **Projects > your\_project > Flow > Analysis**.

<figure><img src="/files/93bW82SCJsegShkI1bWX" alt=""><figcaption></figcaption></figure>

The analysis record will be visible from the Analyses view. The Status will transition through the analysis states as the pipeline progresses. It may take some time (depending on resource availability) for the environment to initialize and the analysis to move to the *In Progress* status. Once the pipeline succeeds, the analysis record will show *Succeeded* as status.

{% hint style="info" %}
This may take considerable time if it is your first analysis due to the required resource management.
{% endhint %}

Once the analysis has succeeded, click the analysis details tab for more information.

<figure><img src="/files/hn0vEX6who1BbIuUPTqe" alt=""><figcaption></figcaption></figure>

From the analysis details view, the logs produced by each process within the pipeline are accessible via the **Steps** tab.

<figure><img src="/files/LOHz0JkovlmGHivcNM2A" alt=""><figcaption></figcaption></figure>

## View Results

Analysis outputs are written to an output folder in the project with the naming convention `{Analysis User Reference}-{Pipeline Code}-{GUID}`. (1)

Inside of the analysis output folder are the files generated by the analysis processes written to the `out` folder. In this tutorial, the file `test.txt` (2) is written to by the `reverse` process. Navigating to the analysis output folder, opening the `test.txt` file details, and selecting the VIEW tab (3) shows the output file contents.

Use the **download** button (4) if you want to download the data to the local machine.

<figure><img src="/files/cBS2F3ALLbzqPnKaIoyQ" alt=""><figcaption></figcaption></figure>


# Nextflow DRAGEN Pipeline

In this tutorial, we will demonstrate how to create and launch a simple DRAGEN pipeline using the Nextflow language in the Platform Core UI. More information about Nextflow on Platform Core can be found [here](/project/p-flow/f-pipelines/pi-nextflow). For this example, we will implement the alignment and variant calling example from this [DRAGEN support page](https://support-docs.illumina.com/SW/DRAGEN_v40/Content/SW/DRAGEN/AligningVariantCallingExamples_fDG_dtREF.htm) for Paired-End FASTQ Inputs.

## Linking a DRAGEN bundle

You need a project in which the pipeline will reside. You can choose an existing project or create a new one. See the [Projects](/home/h-projects) page for information on how to create a project. For this tutorial, we will use a project called *Getting Started*.

Once you have selected or created your project, you need to link a DRAGEN bundle to it give the project access to the DRAGEN docker image. Open your project and navigate to **Projects > your\_project > Project settings > Details > Edit**. From here, select the + symbol next to linked bundles and select a *DRAGEN Demo Tool* bundle to add to the project. For this tutorial, link *DRAGEN Demo Bundle 4.0.3*.

Once the bundle has been linked to your project, you can access the docker image by navigating to the main level and opening **System Settings > Docker Repository.** There, click the docker image *dragen-ica-4.0.3*. At the bottom of the screen, you will see the regions where this bundle is available.

{% hint style="info" %}
The URL presented here will be used later in the `container` directive for your Nextflow DRAGEN process.
{% endhint %}

## Creating the pipeline

Select **Projects > your\_project > Flow > Pipelines**. From the **Pipelines** view, click **+Create > Nextflow > JSON based** to start creating a Nextflow pipeline.

<figure><img src="/files/K0QJSteKOz3hJpkS7HIY" alt="" width="375"><figcaption></figcaption></figure>

### Details

In the Nextflow pipeline creation view, use the **Details** tab to add information about the pipeline. Add values for the required *Code* (pipeline name) and *Description* fields. *Nextflow Version* and *Storage size* defaults to preassigned values.

<figure><img src="/files/MNxClnNKVvzpGnKQEmH1" alt=""><figcaption></figcaption></figure>

### Main.nf

Next, add the Nextflow pipeline definition by navigating to the **Nextflow files > main files > main.nf**. You will see a text editor. Copy and paste the following definition into the text editor. **Modify the `container` directive by replacing the current URL with the URL found in the docker image&#x20;*****dragen-ica-4.0.3*****. (System Settings > Docker Repository > your\_docker\_image > Regions)**.

This pipeline performs the following actions:

1. Accepts one paired FASTQ, one compressed reference file and a sample name
2. Schedules a **FPGA‑backed Kubernetes pod on Platform Core**
3. Unpacks the reference to local scratch
4. Runs DRAGEN with variant calling
5. Uploads **all outputs** to Platform Core cloud storage

```groovy
nextflow.enable.dsl = 2

process DRAGEN {

    // The container must be a DRAGEN image that is included in an accepted bundle and will determine the DRAGEN version
    container '079623148045.dkr.ecr.us-east-1.amazonaws.com/cp-prod/7ecddc68-f08b-4b43-99b6-aee3cbb34524:latest'
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'fpga2-medium'
    pod annotation: 'volumes.illumina.com/scratchSize', value: '1TiB'

    // ICA will upload everything in the "out" folder to cloud storage 
    publishDir 'out', mode: 'symlink'

    input:
        tuple path(read1), path(read2)
        val sample_id
        path ref_tar

    output:
        stdout emit: result
        path '*', emit: output

    script:
        """
        set -ex
        mkdir -p /scratch/reference
        tar -C /scratch/reference -xf ${ref_tar}
        
        /opt/edico/bin/dragen --partial-reconfig HMM --ignore-version-check true
        /opt/edico/bin/dragen --lic-instance-id-location /opt/instance-identity \\
            --output-directory ./ \\
            -1 ${read1} \\
            -2 ${read2} \\
            --intermediate-results-dir /scratch \\
            --output-file-prefix ${sample_id} \\
            --RGID ${sample_id} \\
            --RGSM ${sample_id} \\
            --ref-dir /scratch/reference \\
            --enable-variant-caller true
        """
}

workflow {
    DRAGEN(
        Channel.of([file(params.read1), file(params.read2)]),
        Channel.of(params.sample_id),
        Channel.fromPath(params.ref_tar)
    )
}
```

Refer to the [Nextflow](/project/p-flow/f-pipelines/pi-nextflow) page for details on Platform Core-specific attributes within the Nextflow definition.

* To specify a compute type for a Nextflow process, use the [pod](https://www.nextflow.io/docs/latest/process.html#process-pod) directive within each process.
* Outputs for Nextflow pipelines are uploaded from the `out` folder in the attached shared filesystem. The [publishDir](https://www.nextflow.io/docs/latest/process.html#publishdir) directive specifies the output folder for a given process. Only data moved to the out folder using the `publishDir` directive will be uploaded to the Platform Core project after the pipeline finishes executing.

### Input Form

Next, we create the input form used for the pipeline. This is done on the **Inputform files** tab. More information on the specifications for the input form can be found in [Input Form](/project/p-flow/f-pipelines/pi-inputform) page.

This pipeline takes two FASTQ files, one *reference file* and one *sample\_id* parameter as input.

Paste the following JSON input form into the **inputForm.json** text editor.

```json
{
  "fields": [
    {
      "id": "read1",
      "label": "FASTQ read 1",
      "type": "data",
      "dataFilter": {
        "dataType": "file",
        "dataFormat": ["FASTQ"]
      },
      "maxValues": 1,
      "minValues": 1
    },
    {
      "id": "read2",
      "label": "FASTQ read 2",
      "type": "data",
      "dataFilter": {
        "dataType": "file",
        "dataFormat": ["FASTQ"]
      },
      "maxValues": 1,
      "minValues": 1
    },
    {
      "id": "ref_tar",
      "label": "Reference TAR",
      "type": "data",
      "dataFilter": {
        "dataType": "file",
        "dataFormat": ["TAR"]
      },
      "maxValues": 1,
      "minValues": 1
    },
    {
        "id": "sample_id",
        "type": "textbox",
        "label": "Sample ID"
      }
  ]
}
```

Click the **Simulate** button (bottom left) to preview the launch form fields.

<figure><img src="/files/h5BC0hj1Z0Azvys9TY9w" alt=""><figcaption></figcaption></figure>

Click the `Save` button (top right) to save the changes.

## Running the pipeline

{% hint style="info" %}
If you have no test data available, you need to link the *Dragen Demo Bundle* to your project at **Projects > your\_project > Project Settings > Details > Linked Bundles**.
{% endhint %}

Go to the **projects > your\_project > flow > pipelines > your\_pipeline** and click **Start Analysis**.

Fill in the required fields and click on **Start Analysis** button.

<figure><img src="/files/xSranrejMDCCDhZC4pgy" alt=""><figcaption></figcaption></figure>

#### Results

You can monitor the run from the **Projects > your\_project > Flow > analysis** page. Once the Status changes to Succeeded, you can click on the run to access the results.

## Useful Links

* [Illumina DRAGEN Documentation](https://support-docs.illumina.com/SW/DRAGEN_v40/Content/SW/DRAGEN/GettingStarted_fDG.htm)
* [Official Nextflow documentation](https://www.nextflow.io/)


# Nextflow: Scatter-gather Method

Nextflow supports scatter-gather patterns natively through Channels. The initial [example](/tutorials/nextflow) uses this pattern by splitting the FASTA file into chunks to channel *records* in the task **splitSequences**, then by processing these chunks in the task **reverse**.

In this tutorial, we will create a pipeline which will split a TSV file into chunks, sort them, and merge them together.

## Creating the pipeline

Select **Projects > your\_project > Flow > Pipelines**. From the **Pipelines** view, click the **+Create > Nextflow** **> XML based** button to start creating a Nextflow pipeline.

<figure><img src="/files/3hsF8R1aAlka9H228gWW" alt=""><figcaption></figcaption></figure>

In the **Details** tab, add values for the required *Code* (unique pipeline name) and *Description* fields. *Nextflow Version* and *Storage size* defaults to preassigned values.

First, we present the individual processes. Select **+Nextflow files > + Create** and label the file **split.nf**. Copy and paste the following definition.

```groovy
process split {
    container 'public.ecr.aws/lts/ubuntu:25.10'
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'
    cpus 1
    memory '512 MB'
    
    input:
    path x
    
    output:
    path("split.*.tsv")
    
    """
    split -a10 -d -l3 --numeric-suffixes=1 --additional-suffix .tsv ${x} split.
    """
    }
```

Next, select **+Create** and name the file **sort.nf**. Copy and paste the following definition.

```groovy
process sort {
    container 'public.ecr.aws/lts/ubuntu:25.10'
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'
    cpus 1
    memory '512 MB'
    
    input:
    path x
    
    output:
    path '*.sorted.tsv'
    
    """
    sort -gk1,1 $x > ${x.baseName}.sorted.tsv
    """
}
```

Select **+Create** again and label the file **merge.nf**. Copy and paste the following definition.

```groovy
process merge {
    container 'public.ecr.aws/lts/ubuntu:25.10'
    pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'
    cpus 1
    memory '512 MB'

    publishDir 'out', mode: 'symlink'
    
    input:
    path x
    
    output:
    path 'merged.tsv'
    
    """
    cat $x > merged.tsv
    """
}
```

Add the corresponding main.nf file by navigating to the **Nextflow files > main.nf** tab and copying and pasting the following definition.

```groovy
nextflow.enable.dsl=2
 
include { sort } from './sort.nf'
include { split } from './split.nf'
include { merge } from './merge.nf'
 
 
params.myinput = "test.test"
 
workflow {
    input_ch = Channel.fromPath(params.myinput)
    split(input_ch)
    sort(split.out.flatten())
    merge(sort.out.collect())
}
```

Here, the operators *flatten* and *collect* are used to transform the emitting channels. The *Flatten* operator transforms a channel in such a way that every item of type Collection or Array is flattened so that each single entry is emitted separately by the resulting channel. The collect operator collects all the items emitted by a channel to a List and return the resulting object as a sole emission.

Finally, copy and paste the following XML configuration into the **XML Configuration** tab.

```xml
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition" code="" version="1.0">
    <pd:dataInputs>
        <pd:dataInput code="myinput" format="TSV" type="FILE" required="true" multiValue="false">
            <pd:label>myinput</pd:label>
            <pd:description></pd:description>
        </pd:dataInput>
    </pd:dataInputs>
    <pd:steps/>
</pd:pipeline>
```

Click the Generate button (at the bottom of the text editor) to preview the launch form fields.

Click the **Save** button to save the changes.

## Running the pipeline

Go to the **Pipelines** page from the left navigation pane. Select the pipeline you just created and click **Start New Analysis**.

Fill in the required fields indicated by red "\*" sign and click on **Start** button. You can monitor the run from the **Analyses** page. Once the Status changes to Succeeded, you can click on the run to access the results page.

In **Projects > your\_project > Flow > Analyses > your\_analysis >** **Steps** you can see that the input file is split into multiple chunks, then these chunks are sorted and merged.


# Nextflow CLI

## Nextflow CLI

In this tutorial, we will demonstrate how to create and launch a Nextflow pipeline using the Platform Core command line interface (CLI).

## Installation

Please refer to [these instructions](https://help.ica.illumina.com/command-line-interface/cli-installation) for installing the Platform Core CLI. To authenticate, please follow the steps in the [Authentication](https://help.ica.illumina.com/command-line-interface/cli-authentication) page.

## Tutorial project

In this tutorial, we will create the [Simple RNA-Seq](https://training.nextflow.io/latest/archive/basic_training/rnaseq_pipeline/) pipeline in Platform Core, which includes four processes:

* index creation
* quantification
* FastQC
* MultiQC

We will also upload a Docker container to the Platform Core Docker repository for use within the pipeline.

### main.nf

The 'main.nf' file defines the pipeline that orchestrates various RNASeq analysis processes.

```nf
nextflow.enable.dsl = 2

process INDEX {
   input:
       path transcriptome_file

   output:
       path 'salmon_index'

   script:
       """
       salmon index -t $transcriptome_file -i salmon_index
       """
}

process QUANTIFICATION {
   publishDir 'out', mode: 'symlink'

   input:
       path salmon_index
       tuple path(read1), path(read2)
       val(quant)

   output:
       path "$quant"

   script:
       """
       salmon quant --libType=U -i $salmon_index -1 $read1 -2 $read2 -o $quant
       """
}

process FASTQC {

   input:
       tuple path(read1), path(read2)

   output:
       path "fastqc_logs"

   script:
       """
       mkdir fastqc_logs
       fastqc -o fastqc_logs -f fastq -q ${read1} ${read2}
       """
}

process MULTIQC {
   publishDir 'out', mode:'symlink'

   input:
       path '*'

   output:
       path 'multiqc_report.html'

   script:
       """
       multiqc .
       """
}

workflow {
   index_ch = INDEX(Channel.fromPath(params.transcriptome_file))
   quant_ch = QUANTIFICATION(index_ch, Channel.of([file(params.read1), file(params.read2)]),Channel.of("quant"))
   fastqc_ch = FASTQC(Channel.of([file(params.read1), file(params.read2)]))
   MULTIQC(quant_ch.mix(fastqc_ch).collect())
}
```

The script uses the following tools:

* **Salmon**: Software tool for quantification of transcript abundance from RNA-seq data.
* **FastQC**: QC tool for sequencing data
* **MultiQC**: Tool to aggregate and summarize QC reports

We need a Docker container containing these tools. For the sake of this tutorial, we will use the container from the original tutorial. You can refer to the ["Build and push your own Docker image to Platform Core"](/tutorials/cli-cwl/cwl-graphical-pipeline#build-and-push-your-own-docker-image-to-platform-core) section to build your own docker image with the required tools.

## Docker image upload

With [Docker installed](https://docs.docker.com/desktop/) in your computer, download the image required for this project using the following command.

`docker pull nextflow/rnaseq-nf`

Create a tarball of the image to upload to Platform Core.

```
docker save nextflow/rnaseq-nf > cont_rnaseq.tar
```

Following are lists of commands that you can use to upload the tarball to your project.

```
# Enter the project context
icav2 enter docs
# Upload the container image to the root directory (/) of the project
icav2 projectdata upload cont_rnaseq.tar /
```

**Add the image to the Platform Core Docker repository**

The uploaded image can be added to the Platform Core Docker repository from the graphical user interface.

Change the format for the image tarball to DOCKER:

1. Navigate to **Projects > your\_project > Data**.
2. Check the checkbox for the uploaded tarball.
3. Click on **Manage > Change format**.
4. In the new popup window, select "DOCKER" format and save.

To add this image to the Platform Core Docker repository, first click on **Projects** to go back to the home page.

1. From the Platform Core home page, click on **System Settings > Docker Repository > Create > Image**.
2. This will open a new window that lets you select the region (US, EU, CA) in which your your project is and the docker image from the bottom pane.
3. Edit the Name field to rename it. For this tutorial, we will change the name to "rnaseq". Select the region, and give it a version number, and description. Click on "Save".

{% hint style="info" %}
If you have the images hosted in other repositories, you can add them as external image by using **System** **Settings > Docker Repository > Create > External Image**.
{% endhint %}

After creating a new docker image, you can click on the image to get the container URL (under Regions) for the nextflow configuration file.

#### Nextflow configuration file

Create a configuration file called "nextflow\.config" in the same folder as the main.nf file above. Use the URL copied above to add the `process.container` line in the config file.

{% code overflow="wrap" %}

```
process.container = '079623148045.dkr.ecr.us-east-1.amazonaws.com/cp-prod/3cddfc3d-2431-4a85-82bb-dae061f7b65d:latest'
```

{% endcode %}

You can add a pod directive within a process or in the config file to specify a compute type. The following is an example of a configuration file with the 'standard-small' compute type for all processes. Please refer to the [Compute Types](/project/p-flow/f-pipelines#compute-resources) page for a list of available compute types.

{% code overflow="wrap" %}

```
process {
    container = '079623148045.dkr.ecr.us-east-1.amazonaws.com/cp-prod/3cddfc3d-2431-4a85-82bb-dae061f7b65d:latest'
    pod = [
        annotation: 'scheduler.illumina.com/presetSize',
        value: 'standard-small'
    ]  
}
```

{% endcode %}

#### Parameters file

The parameters file defines the pipeline input parameters. Refer to the [JSON](/project/p-flow/f-pipelines/json-based-input-forms) or [XML](/project/p-flow/f-pipelines/pi-inputform) input for detailed information for creating correctly formatted parameters files.

An empty form looks as follows:

{% code overflow="wrap" %}

```
<pipeline code="" version="1.0" xmlns="xsd://www.illumina.com/ica/cp/pipelinedefinition">
   <dataInputs>
   </dataInputs>
   <steps>
   </steps>
</pipeline>
```

{% endcode %}

The input files are specified within a single **dataInputs** node with individual input file specified in a separate **dataInput** node. Settings (as opposed to files) are specified within the **steps** node. Settings represent any non-file input to the pipeline, including but not limited to, strings, booleans, integers, etc..

For this tutorial, we do not have any settings parameters but it requires multiple file inputs. The parameters.xml file looks as follows:

{% code overflow="wrap" %}

```
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition" code="" version="1.0">
   <pd:dataInputs>
       <pd:dataInput code="read1" format="FASTQ" type="FILE" required="true" multiValue="false">
           <pd:label>FASTQ Read 1</pd:label>
           <pd:description>FASTQ Read 1</pd:description>
       </pd:dataInput>
       <pd:dataInput code="read2" format="FASTQ" type="FILE" required="true" multiValue="false">
           <pd:label>FASTQ Read 2</pd:label>
           <pd:description>FASTQ Read 2</pd:description>
       </pd:dataInput>
       <pd:dataInput code="transcriptome_file" format="FASTA" type="FILE" required="true" multiValue="false">
           <pd:label>Transcript</pd:label>
           <pd:description>Transcript faster</pd:description>
       </pd:dataInput>
   </pd:dataInputs>
   <pd:steps/>
</pd:pipeline>
```

{% endcode %}

Use the following commands to create the pipeline with the above contents in your project.

If not already in the project context, enter it by using the following command:

`icav2 enter <PROJECT NAME or ID>`

Create pipeline using `icav2 project pipelines create nextflow` Example:

{% code overflow="wrap" %}

```
icav2 projectpipelines create nextflow rnaseq-docs --main main.nf --parameter parameters.xml --config nextflow.config --storage-size small --description 'cli nextflow pipeline'
```

{% endcode %}

If you prefer to organize the processes in different folders/files, you can use `--other` parameter to upload the different processes as additional files. Example:

{% code overflow="wrap" %}

```
icav2 projectpipelines create nextflow rnaseq-docs --main main.nf --parameter parameters.xml --config nextflow.config --other index.nf:filename=processes/index.nf --other quantification.nf:filename=processes/quantification.nf --other fastqc.nf:filename=processes/fastqc.nf --other multiqc.nf:filename=processes/multiqc.nf --storage-size small --description 'cli nextflow pipeline'
```

{% endcode %}

You can refer to [Nextflow: Pipeline Lift](broken://pages/rQNZoMf1zM8oyNSOzOc5) page to explore options to automate this process.

Refer to [Launch Pipelines on CLI](broken://pages/1zoggcjo24Y2HsjoINyk) for details on running the pipeline from CLI.

Example command to run the pipeline from CLI:

{% code overflow="wrap" %}

```
icav2 projectpipelines start nextflow <pipeline_id> --input read1:<read1_file_id> --input read2:<read2_file_id> --input transcriptome_file:<transcriptome_file_id> --storage-size small --user-reference demo_run
```

{% endcode %}

You can get the pipeline id under "ID" column by running the following command:

```
icav2 projectpipelines list
```

You can get the file ids under "ID" column by running the following commands:

```
icav2 projectdata list
```

Please refer to command help (`icav2 [command] --help`) to determine available flags to filter output of above commands if necessary. You can also refer to [Command Index](/command-line-interface/cli-indexcommands) page for available flags for the icav2 commands.

For more help on uploading data to Platform Core, please refer to the [Data Transfer options](/command-line-interface/cli-datatransfer) page.


# CWL CLI Pipeline Execution

In this tutorial, we will demonstrate how to create and launch a pipeline using the CWL language with the Platform Core command line interface (CLI).

## Installation

Please refer to [these instructions](/command-line-interface/cli-installation) for installing the Platform Core CLI.

## Tutorial project

In this project, we will create two simple tools and build a pipeline that we can run on Platform Core using the CLI. The first tool (tool-fqTOfa.cwl) will convert a FASTQ file to a FASTA file. The second tool(tool-countLines.cwl) will count the number of lines in an input FASTA file. The workflow\.cwl will combine the two tools to convert an input FASTQ file to a FASTA file and count the number of lines in the resulting FASTA file.

Following are the two CWL tools and scripts we will use in the project. If you are new to CWL, please refer to the cwl [user guide](https://www.commonwl.org/user_guide/) for a better understanding of CWL codes. You will also need the cwltool installed to create these tools and processes. You can find installation instructions on the CWL [github](https://github.com/common-workflow-language/cwltool) page.

### tool-fqTOfa.cwl

```{cwl}
#!/usr/bin/env cwltool

cwlVersion: v1.0
class: CommandLineTool
inputs:
  inputFastq:
    type: File
    inputBinding:
        position: 1
stdout: test.fasta
outputs:
  outputFasta:
    type: File
    streamable: true
    outputBinding:
        glob: test.fasta

arguments:
- 'NR%4 == 1 {print ">" substr($0, 2)}NR%4 == 2 {print}'
baseCommand:
- awk
```

### tool-countLines.cwl

```{cwl}
#!/usr/bin/env cwltool

cwlVersion: v1.0
class: CommandLineTool
baseCommand: [wc, -l]
inputs:
  inputFasta:
    type: File
    inputBinding:
        position: 1
stdout: lineCount.tsv
outputs:
  outputCount:
    type: File
    streamable: true
    outputBinding:
        glob: lineCount.tsv
```

### workflow\.cwl

```{cwl}
cwlVersion: v1.0
class: Workflow
inputs:
  ipFQ: File

outputs:
  count_out:
    type: File
    outputSource: count/outputCount
  fqTOfaOut:
    type: File
    outputSource: convert/outputFasta
   
steps:
  convert:
    run: tool-fqTOfa.cwl
    in:
      inputFastq: ipFQ
    out: [outputFasta]

  count:
    run: tool-countLines.cwl
    in:
      inputFasta: convert/outputFasta
    out: [outputCount]
```

{% hint style="warning" %}
we don't specify the Docker image used in both tools. In such a case, the default behaviour is to use public.ecr.aws/docker/library/bash:5 image. This image contains basic functionality (sufficient to execute `wc` and `awk` commands).
{% endhint %}

If you want to use a different public image, you can specify it using *requirements* tag in cwl file. Assuming you want to use \*ubuntu:latest' you need to add

```{cwl}
requirements:
  - class: DockerRequirement
    dockerPull: ubuntu:latest
```

If you want to use a Docker image from the Platform Core Docker repository, you need the link to AWS ECR from the Platform Core GUI. Double-click on the image name in the Docker repository and copy the URL to the clipboard. Add the URL to *dockerPull* key.

```{cwl}
requirements:
  - class: DockerRequirement
    dockerPull: 079623148045.dkr.ecr.eu-central-1.amazonaws.com/cp-prod/XXXXXXXXXX:latest
```

To add a custom or public docker image to the Platform Core repository, refer to the [Docker Repository](https://help.ica.illumina.com/home/h-dockerrepository).

## Authentication

Before you can use the Platform Core CLI, you need to authenticate using the Illumina API key. Follow [these instructions](https://help.ica.illumina.com/command-line-interface/cli-authentication) to authenticate.

## Enter/Create a Project

Either create a project or use an existing project to create a new pipeline. You can create a new project using the `icav2 projects create` command.

```{bash}
% icav2 projects create basic-cli-tutorial --region c39b1feb-3e94-4440-805e-45e0c76462bf
```

If you do not provide the -`-region` flag, the value defaults to the existing region when there is only one region available. When there is more than one region available, a selection must be made from the available regions at the command prompt. The region input can be determined by calling the `icav2 regions list` command first.

You can select the project to work on by entering the project using the `icav2 projects enter` command. Thus, you won't need to specify the project as an argument.

```{bash}
% icav2 projects enter basic-cli-tutorial
```

You can also use the `icav2 projects list` command to determine the names and ids of the project you have access to.

```{bash}
% icav2 projects list
```

## Create a pipeline on Platform Core

`projectpipelines` is the root command to perform actions on pipelines in a project. The `create` command creates a pipeline in the current project.

The parameter file specifies the input with additional parameter settings for each step in the pipeline. In this tutorial, input is a FASTQ file shown inside \<dataInput> tag in the parameter file. There aren't any specific settings for the pipeline steps resulting in a parameter file below with an empty \<steps> tag. Create a parameter file (parameters.xml) with the following content using a text editor.

```{xml}
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition" code="" version="1.0">
    <pd:dataInputs>
        <pd:dataInput code="ipFQ" format="FASTQ" type="FILE" required="true" multiValue="false">
            <pd:label>ipFQ</pd:label>
            <pd:description></pd:description>
        </pd:dataInput>
    </pd:dataInputs>
    <pd:steps/>
</pd:pipeline>
```

The following command creates a pipeline called "cli-tutorial" using the workflow\.cwl, tools "tool-fqTOfa.cwl" and "tool-countLines.cwl" and parameter file "parameter.xml" with small storage size.

```{bash}
% icav2 projectpipelines create cwl cli-tutorial --workflow workflow.cwl --tool tool-fqTOfa.cwl --tool tool-countLines.cwl --parameter parameters.xml --storage-size small --description "cli tutorial pipeline"
```

Once the pipeline is created, you can view it using the `list` command.

```{bash}
% icav2 projectpipelines list
ID                                   	CODE                      	DESCRIPTION                                      
6779fa3b-e2bc-42cb-8396-32acee8b6338	cli-tutorial             	cli tutorial pipeline 
```

## Running the pipeline

Upload data to the project using the `icav2 projectdata upload` command. Refer to the [Data page](https://help.ica.illumina.com/project/p-data) for advanced data upload features. For this test, we will use a small FASTQ file test.fastq containing the following reads.

```{txt}
@SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
AAGTTACCCTTAACAACTTAAGGGTTTTCAAATAGA
+SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
IIIIIIIIIIIIIIIIIIIIDIIIIIII>IIIIII/
@SRR001666.2 071112_SLXA-EAS1_s_7:5:1:801:338 length=36
AGCAGAAGTCGATGATAATACGCGTCGTTTTATCAT
+SRR001666.2 071112_SLXA-EAS1_s_7:5:1:801:338 length=36
IIIIIIIIIIIIIIIIIIIIIIGII>IIIII-I)8I
@SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
AAGTTACCCTTAACAACTTAAGGGTTTTCAAATAGA
+SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
IIIIIIIIIIIIIIIIIIIIDIIIIIII>IIIIII/
@SRR001666.2 071112_SLXA-EAS1_s_7:5:1:801:338 length=36
AGCAGAAGTCGATGATAATACGCGTCGTTTTATCAT
+SRR001666.2 071112_SLXA-EAS1_s_7:5:1:801:338 length=36
IIIIIIIIIIIIIIIIIIIIIIGII>IIIII-I)8I
@SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
AAGTTACCCTTAACAACTTAAGGGTTTTCAAATAGA
+SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
IIIIIIIIIIIIIIIIIIIIDIIIIIII>IIIIII/
@SRR001666.2 071112_SLXA-EAS1_s_7:5:1:801:338 length=36
AGCAGAAGTCGATGATAATACGCGTCGTTTTATCAT
+SRR001666.2 071112_SLXA-EAS1_s_7:5:1:801:338 length=36
IIIIIIIIIIIIIIIIIIIIIIGII>IIIII-I)8I
```

The `icav2 projectdata upload` command lets you upload data to Platform Core.

```{bash}
% icav2 projectdata upload test.fastq /
oldFilename= test.fastq en newFilename= test.fastq
bucket= stratus-gds-use1  prefix= 0a488bb2-578b-404a-e09d-08d9e3343b2b/test.fastq
Using: 1 workers to upload 1 files
15:23:32: [0]  Uploading /Users/user1/Documents/icav2_validation/for_tutorial/working/test.fastq
15:23:33: [0]  Uploaded /Users/user1/Documents/icav2_validation/for_tutorial/working/test.fastq to /test.fastq in 794.511591ms
Finished uploading 1 files in 795.244677ms

```

The `list` command lets you view the uploaded file. Note the ID of the file you want to use with the pipeline.

```{bash}
% icav2 projectdata list                
PATH          NAME        TYPE  STATUS    ID                                    OWNER                                 
/test.fastq  test.fastq FILE  AVAILABLE fil.c23246bd7692499724fe08da020b1014  4b197387-e692-4a78-9304-c7f73ad75e44
```

The `icav2 projectpipelines start` command initiates the pipeline run. The following command runs the pipeline. Write down the id for exploring the analysis later.

If for some reason your `create` command fails and needs to rerun, you might get an error (ConstraintViolationException). If so, try your command with a different name.

```{bash}
% icav2 projectpipelines start cwl cli-tutorial --type-input STRUCTURED --input ipFQ:fil.c23246bd7692499724fe08da020b1014 --user-reference tut-test
analysisStorage.description           1.2 TB
analysisStorage.id                    6e1b6c8f-f913-48b2-9bd0-7fc13eda0fd0
analysisStorage.name                  Small
analysisStorage.ownerId               8ec463f6-1acb-341b-b321-043c39d8716a
analysisStorage.tenantId              f91bb1a0-c55f-4bce-8014-b2e60c0ec7d3
analysisStorage.tenantName            ica-cp-admin
analysisStorage.timeCreated           2021-11-05T10:28:20Z
analysisStorage.timeModified          2021-11-05T10:28:20Z
id                                    461d3924-52a8-45ef-ab62-8b2a29621021
ownerId                               7fa2b641-1db4-3f81-866a-8003aa9e0818
pipeline.analysisStorage.description  1.2 TB
pipeline.analysisStorage.id           6e1b6c8f-f913-48b2-9bd0-7fc13eda0fd0
pipeline.analysisStorage.name         Small
pipeline.analysisStorage.ownerId      8ec463f6-1acb-341b-b321-043c39d8716a
pipeline.analysisStorage.tenantId     f91bb1a0-c55f-4bce-8014-b2e60c0ec7d3
pipeline.analysisStorage.tenantName   ica-cp-admin
pipeline.analysisStorage.timeCreated  2021-11-05T10:28:20Z
pipeline.analysisStorage.timeModified 2021-11-05T10:28:20Z
pipeline.code                         cli-tutorial
pipeline.description                  Test, prepared parameters file from working GUI
pipeline.id                           6779fa3b-e2bc-42cb-8396-32acee8b6338
pipeline.language                     CWL
pipeline.ownerId                      7fa2b641-1db4-3f81-866a-8003aa9e0818
pipeline.tenantId                     d0696494-6a7b-4c81-804d-87bda2d47279
pipeline.tenantName                   icav2-entprod
pipeline.timeCreated                  2022-03-10T13:13:05Z
pipeline.timeModified                 2022-03-10T13:13:05Z
reference                             tut-test-cli-tutorial-eda7ee7a-8c65-4c0f-bed4-f6c2d21119e6
status                                REQUESTED
summary                               
tenantId                              d0696494-6a7b-4c81-804d-87bda2d47279
tenantName                            icav2-entprod
timeCreated                           2022-03-10T20:42:42Z
timeModified                          2022-03-10T20:42:43Z
userReference                         tut-test
```

You can check the status of the run using the `icav2 projectanalyses get` command.

```{bash}
%   icav2 projectanalyses get 461d3924-52a8-45ef-ab62-8b2a29621021
analysisStorage.description           1.2 TB
analysisStorage.id                    6e1b6c8f-f913-48b2-9bd0-7fc13eda0fd0
analysisStorage.name                  Small
analysisStorage.ownerId               8ec463f6-1acb-341b-b321-043c39d8716a
analysisStorage.tenantId              f91bb1a0-c55f-4bce-8014-b2e60c0ec7d3
analysisStorage.tenantName            ica-cp-admin
analysisStorage.timeCreated           2021-11-05T10:28:20Z
analysisStorage.timeModified          2021-11-05T10:28:20Z
endDate                               2022-03-10T21:00:33Z
id                                    461d3924-52a8-45ef-ab62-8b2a29621021
ownerId                               7fa2b641-1db4-3f81-866a-8003aa9e0818
pipeline.analysisStorage.description  1.2 TB
pipeline.analysisStorage.id           6e1b6c8f-f913-48b2-9bd0-7fc13eda0fd0
pipeline.analysisStorage.name         Small
pipeline.analysisStorage.ownerId      8ec463f6-1acb-341b-b321-043c39d8716a
pipeline.analysisStorage.tenantId     f91bb1a0-c55f-4bce-8014-b2e60c0ec7d3
pipeline.analysisStorage.tenantName   ica-cp-admin
pipeline.analysisStorage.timeCreated  2021-11-05T10:28:20Z
pipeline.analysisStorage.timeModified 2021-11-05T10:28:20Z
pipeline.code                         cli-tutorial
pipeline.description                  Test, prepared parameters file from working GUI
pipeline.id                           6779fa3b-e2bc-42cb-8396-32acee8b6338
pipeline.language                     CWL
pipeline.ownerId                      7fa2b641-1db4-3f81-866a-8003aa9e0818
pipeline.tenantId                     d0696494-6a7b-4c81-804d-87bda2d47279
pipeline.tenantName                   icav2-entprod
pipeline.timeCreated                  2022-03-10T13:13:05Z
pipeline.timeModified                 2022-03-10T13:13:05Z
reference                             tut-test-cli-tutorial-eda7ee7a-8c65-4c0f-bed4-f6c2d21119e6
startDate                             2022-03-10T20:42:42Z
status                                SUCCEEDED
summary                               
tenantId                              d0696494-6a7b-4c81-804d-87bda2d47279
tenantName                            icav2-entprod
timeCreated                           2022-03-10T20:42:42Z
timeModified                          2022-03-10T21:00:33Z
userReference                         tut-test
```

The pipelines can be run using JSON input type as well. The following is an example of running pipelines using JSON input type.

{% hint style="info" %}
JSON input works only with file-based CWL pipelines built using code, not a graphical editor in Platform Core.
{% endhint %}

```{bash}
 % icav2 projectpipelines start cwl cli-tutorial --data-id fil.c23246bd7692499724fe08da020b1014 --input-json '{
  "ipFQ": {
    "class": "File",
    "path": "test.fastq"
  }
}' --type-input JSON --user-reference tut-test-json
```

## Notes

### runtime.ram and runtime.cpu

`runtime.ram` and `runtime.cpu` values are by default evaluated using the compute environment running the host CWL runner. CommandLineTool Steps within a CWL pipeline run on different compute environments than the host CWL runner, so the valuations of the `runtime.ram` and `runtime.cpu` for within the CommandLineTool will not match the runtime environment the tool is running in. The valuation of `runtime.ram` and `runtime.cpu` can be overridden by specifying `cpuMin` and `ramMin` in the `ResourceRequirements` for the CommandLineTool.


# CWL Graphical Pipeline

This tutorial aims to guide you through the process of creating CWL tools and pipelines from the very beginning. By following the steps and techniques presented here, you will gain the necessary knowledge and skills to develop your own pipelines or transition existing ones to Platform Core.

## Build and push your own Docker image to Platform Core

The foundation for every tool in Platform Core is a Docker image (externally published or created by the user). Here we present how to create your own Docker image for the popular tool (FASTQC).

Copy the contents displayed below to a text editor and save it as a Dockerfile. Make sure you use an editor which does not add formatting to the file.

```dockerfile
FROM centos:7
WORKDIR /usr/local

# DEPENDENCIES
RUN yum -y install java-1.8.0-openjdk wget unzip perl && \
    yum clean all && \
    rm -rf /var/cache/yum

# INSTALLATION fastqc
RUN wget http://www.bioinformatics.babraham.ac.uk/projects/fastqc/fastqc_v0.11.9.zip --no-check-certificate && \
    unzip fastqc_v0.11.9.zip && \
    chmod a+rx /usr/local/FastQC/fastqc && rm -rf fastqc_v0.11.9.zip

# Adding FastQC to the PATH
ENV PATH $PATH:/usr/local/FastQC

# DEFAULTS
ENV LANG=en_US.UTF-8
ENV LC_ALL=en_US.UTF-8
ENTRYPOINT []

## how to build the docker image
## docker build --file fastqc-0.11.9.Dockerfile --tag fastqc-0.11.9:0 .
## docker run --rm -i -t --entrypoint /bin/bash fastqc-0.11.9:0
```

Open a terminal window, place this file in a dedicated folder and navigate to this folder location. Then use the following command:

```
docker build --file fastqc-0.11.9.Dockerfile --tag fastqc-0.11.9:1 .
```

Check the image has been successfully built:

```
docker images
```

Check that the container is functional:

```
docker run --rm -i -t --entrypoint /bin/bash fastqc-0.11.9:1
```

Once inside the container check that the **`fastqc`** command is responsive and prints the expected help message. Remember to **`exit`** the container.

Save a tar of the previously built image locally:

```
docker save fastqc-0.11.9:1 -o fastqc-0.11.9:1.tar.gz
```

Upload your docker image `.tar` to a Platform Core project using browser upload, Connector, or CLI.

{% hint style="warning" %}
In **Projects > your\_project > Data**, select the uploaded .tar file, then click **Manage > Change Format** , select DOCKER and Save.
{% endhint %}

Now go outside of the Project and go to System **Settings > Docker Repository**, Select **Create > Image**. Select your docker file and fill out a name and version and set your type to tool and Press Select.

## Create a CWL tool

While outside of any Project go to **System Settings > Tool Repository** and Select **+Create**. Fill the mandatory fields (Name and Version) and look for a Docker image to link to the tool.

Tool creation in Platform Core adheres to the [CWL standard](https://www.commonwl.org/v1.0/).

You can create a tool by either pasting the tool definition in the code syntax field on the right or you can use the different tabs to manually define inputs, outputs, arguments, settings, etc …

In this tutorial we will use the CWL tool syntax method. Paste the following content in the General tab.

{% hint style="info" %}
Other tabs, except for the Details tab can also be used.
{% endhint %}

```cwl
#!/usr/bin/env cwl-runner

# (Re)generated by BlueBee Platform

$namespaces:
  ilmn-tes: http://platform.illumina.com/rdf/iap/
cwlVersion: cwl:v1.0
class: CommandLineTool
label: FastQC
doc: FastQC aims to provide a simple way to do some quality control checks on raw
  sequence data coming from high throughput sequencing pipelines.
inputs:
  Fastq1:
    type: File
    inputBinding:
      position: 1
  Fastq2:
    type:
    - File
    - 'null'
    inputBinding:
      position: 3
outputs:
  HTML:
    type:
      type: array
      items: File
    outputBinding:
      glob:
      - '*.html'
  Zip:
    type:
      type: array
      items: File
    outputBinding:
      glob:
      - '*.zip'
arguments:
- position: 4
  prefix: -o
  valueFrom: $(runtime.outdir)
- position: 1
  prefix: -t
  valueFrom: '2'
baseCommand:
- fastqc
```

Since the user needs to specify the output folder for FASTQC application (*-o* prefix), we are using the *$(runtime.outdir)* runtime parameter to point to the designated output folder.

## Create the pipeline

Navigate to **Projects > your\_project > Flow > Pipelines > +Create > CWL Graphical**.

Fill the mandatory fields and click on the Definition tab to open the Graphical Editor.

Expand the **Tool Repository** menu (lower right) and drag your FastQC tool into the Editor field (center).

Now drag one Input and one Output file icon (on top) into the Editor field as well. Both may be given a Name (editable fields on the right when icon is selected) and need a Format attribute. Set the Input Format to fastq and Output Format to html. Connect both Input and Output files to the matching nodes on the tool itself (mouse over the node, then hold-click and drag to connect).

Press Save, you just created your first **FastQC** pipeline on Platform Core!

![FastQC](/files/3T6aVMm1GnJKG2DVgctA)

## Run a pipeline

First make sure you have at least one Fastq file uploaded and/or linked to your Project. You may use Fastq files available in the Bundle.

Navigate to Pipelines and select the pipeline you just created, then press **Start analysis**

Fill the mandatory fields and click on the **+** button to open the File Selection dialog box. Select one of the Fastq files available to you.

Press **Start analysis** on the top right, the platform is now orchestrating the pipeline execution.

## View Results

Navigate to **Projects > your\_project > Flow > Analyses** and observe that the pipeline execution is now listed and will first be in Status *Requested*. After a few minutes the Status should change to *In Progress* and then to *Succeeded*.

Once this Analysis succeeds click it to enter the Analysis details view. You will see the FastQC HTML output file listed on the Output files tab. Click on the file to open Data Details view. Since it is an HTML file Format there is a View tab that allows visualizing the HTML within the browser.


# CWL DRAGEN Pipeline

In this tutorial, we will demonstrate how to create and launch a DRAGEN pipeline using the CWL language.

In Platform Core, CWL pipelines are built using tools developed in CWL. For this tutorial, we will use the "DRAGEN Demo Tool" included with DRAGEN Demo Bundle 3.9.5.

## Linking bundle to Project

1. Start by selecting a project at the **Projects** inventory.

<img src="/files/yHtJppPeUr3w0W8BtkQi" alt="" width="316">

2. In the details page, select `Edit`.

![](/files/IVxHClkliGWk7IpZIXFh)

3. In the *edit mode of the details page*, click the `+` button in the LINKED BUNDLES section.

![](/files/tjBuXx5TWHTuY6qLIkWF)

4. In the *Add Bundle to Project* window:\
   Select the DRAGEN demo tool bundle from the list. Once you have selected the bundle, the **Link Bundles** button becomes available. Select it to continue.

{% hint style="info" %}
You can select multiple bundles using `Ctrl + Left mouse button` or `Shift + Left mouse button`.
{% endhint %}

![](/files/US3Zy43SdzzTmqhpQMf7)

5. In the project details page, the selected bundle will appear under the LINKED BUNDLES section. If you need to remove a bundle, click on the `-` button. Click **Save** to save the project with linked bundles.

![](/files/lvVQgC5u8uxfutvs8ydV)

## Create Pipeline

1. From the project details page, select **Pipelines > CWL**

![](/files/s0D2ZUgMwlCgWqLzcTZi)

2. You will be given options to create pipelines using a graphical interface or code. For this tutorial, we will select Graphical.

![](/files/Igo1Xpb7uLjJSYvbesyY)

3. Once you have selected the Graphical option, you will see a page with multiple tabs. The first tab is the *Information* page where you enter pipeline information. You can find the details for different fields in the tab in the [GitBook](https://help.ica.illumina.com/project/p-flow/f-pipelines). The following three fields are required for the INFORMATION page.
   * **Code**: Provide pipeline name here.
   * **Description**: Provide pipeline description here.
   * **Storage size**: Select the storage size from the drop-down menu.

![](/files/NWyMCw3wWYrTgWoaVNEc)

4. The *Documentation* tab provides options for configuring the HTML description for the tool. The description appears in the tool repository but is excluded from exported CWL definitions.
5. The *Definition* tab is used to define the pipeline. When using graphical mode for the pipeline definition, the Definition tab provides options for configuring the pipeline using a visualization panel (A) and a list of component menus (B). You can find details on each section in the component menu [here](https://help.ica.illumina.com/project/p-flow/f-pipelines#definition)

<figure><img src="/files/d6B0Q22I2zyhDmDgUFgV" alt=""><figcaption></figcaption></figure>

6. To build a pipeline, start by selecting **Machine PROFILE** from the component menu section on the right. All fields are required and are pre-filled with default values. Change them as needed.
   * The profile ***Name*** field will be updated based on the selected Resource. You can change it as needed.
   * ***Color*** assigns the selected color to the tool in the design view to easily identify the machine profile when more than one tool is used in the pipeline.
   * ***Tier*** lets you select Standard or Economy tier for AWS instances. Standard is on-demand ec2 instance and Economy is spot ec2 instance. You can find the difference between the two AWS instances [here](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-spot-instances.html). You can find the price difference between the two tiers [here](https://help.ica.illumina.com/reference/r-pricing#compute).
   * ***Resource*** lets you choose from various compute resources available. In this case, we are building a DRAGEN pipeline and we will need to select a resource with FPGA in it. Choose from FPGA resources (FPGA Medium/Large) based on your needs.

![](/files/DS3UepimHEoCSzKqiBAX)

7. Once you have selected the Machine Profile for the tool, find your tool from the *Tool Repository* at the bottom section of the component menu on the right. In this case, we are using the *DRAGEN Demo Tool*. Drag and drop the tool from the Tool Repository section to the visualization panel.

![](/files/f5cUk5jSDZ7aBuimLSDA)

8. The dropped tool will show the machine profile color, number of outputs and inputs, and warning to indicate missing parameters, mandatory values, and connections. Selecting the tool in the visualization panel activates the tool (DRAGEN Demo Tool) component menu. On the component menu section, you will find the details of the tool under *Tool - DRAGEN Demo Tool*. This section lists the inputs, outputs, additional parameters, and the machine profile required for the tool. In this case, the DRAGEN Demo Tool requires three inputs (FASTQ read 1, FASTQ read 2, and a Reference genome). The tool has two outputs (a VCF file and an output folder). The tool also has a mandatory parameter (Output File Prefix). Enter the value for the input parameter (Output File Prefix) in the text box.

![](/files/PaHnZhKKO3G4UcSyrKSH)

9. The top right corner of the visualization panel has icons to zoom in and out in the visualization panel followed by three icons: ref, in, and out. Based on the type of input/output needed, drag and drop the icons into the visualization area. In this case, we need three inputs (read 1, read 2, and Reference hash table.) and two outputs (VCF file and output folder). Start by dragging and dropping the first input (a). Connect the input to the tool by clicking on the blue dot at the bottom of the input icon and dragging it to the blue dot representing the first input on the tool (b). Select the input icon to activate the input component menu. The input section for the first input lets you enter the Name, Format, and other relevant information based on tool requirements. In this case, for the first input, enter the following information:
   * Name: FASTQ read 1
   * Format: FASTQ
   * Comments: any optional comments

![](/files/4HHANLgAumNPG8hePvnj)

10. Repeat the step for other inputs. Note that the Reference hash table is treated as the input for the tool rather than *Reference files*. So, use the input icon instead of the reference icon.
11. Repeat the process for two outputs by dragging and connecting them to the tool. Note that when connecting output to the tool, you will need to click on the blue dot at the bottom of the tool and drag it to the output.

![](/files/tdFmbnNN5eYJugUCw3h1)

12. Select the tool and enter additional parameters. In this case, the tool requires *Output File Prefix*. Enter *demo*\_ in the text box.
13. Click on the Save button to save the pipeline. Once saved, you can run it from the Pipelines page under Flow from the left menus as any other pipeline.


# CWL: Scatter-gather Method

In bioinformatics and computational biology, the vast and growing amount of data necessitates methods and tools that can process and analyze data in parallel. This demand gave birth to the **scatter-gather** approach, an essential pattern in creating pipelines that offers efficient data handling and parallel processing capabilities. In this tutorial, we will demonstrate how to create a CWL pipeline utilizing the scatter-gather approach. To this purpose, we will use two widely known tools: [fastp](https://academic.oup.com/bioinformatics/article/34/17/i884/5093234) and [multiqc](https://multiqc.info/). Given the functionalities of both fastp and multiqc, their combination in a scatter-gather pipeline is incredibly useful. Individual datasets can be scattered across resources for parallel preprocessing with fastp. Subsequently, the outputs from each of these parallel tasks can be gathered and fed into multiqc, generating a consolidated quality report. This method not only accelerates the preprocessing of large datasets but also offers an aggregated perspective on data quality, ensuring that subsequent analyses are built upon a robust foundation.

## Creating the tools

First, we create the two tools: fastp and multiqc. For this, we need the corresponding Docker images and CWL tool definitions. Please, look up this [part](https://help.ica.illumina.com/home/h-toolrepository#import-tool) of our help sites to learn more how to import a tool into Platform Core. In a nutshell, once the CWL tool definition is pasted into the editor, the other tabs for editing the tool will be populated. To complete the tool, the user needs to select the corresponding Docker image and to provide a tool version (could be any string).

For this demo, we will use the publicly available Docker images: quay.io/biocontainers/fastp:0.20.0--hdbcaa40\_0 for fastp and docker.io/ewels/multiqc:v1.15 for multiqc. In this [tutorial](https://help.ica.illumina.com/home/h-dockerrepository) one can find how to import publicly available Docker images into Platform Core.

Furthermore, we will use the following CWL tool definitions:

```yaml
#!/usr/bin/env cwl-runner

cwlVersion: v1.0
class: CommandLineTool
requirements:
- class: InlineJavascriptRequirement
label: fastp
doc: Modified from https://github.com/nigyta/bact_genome/blob/master/cwl/tool/fastp/fastp.cwl
inputs:
  fastq1:
    type: File
    inputBinding:
      prefix: -i
  fastq2:
    type:
    - File
    - 'null'
    inputBinding:
      prefix: -I
  threads:
    type:
    - int
    - 'null'
    default: 1
    inputBinding:
      prefix: --thread
  qualified_phred_quality:
    type:
    - int
    - 'null'
    default: 20
    inputBinding:
      prefix: --qualified_quality_phred
  unqualified_phred_quality:
    type:
    - int
    - 'null'
    default: 20
    inputBinding:
      prefix: --unqualified_percent_limit
  min_length_required:
    type:
    - int
    - 'null'
    default: 50
    inputBinding:
      prefix: --length_required
  force_polyg_tail_trimming:
    type:
    - boolean
    - 'null'
    inputBinding:
      prefix: --trim_poly_g
  disable_trim_poly_g:
    type:
    - boolean
    - 'null'
    default: true
    inputBinding:
      prefix: --disable_trim_poly_g
  base_correction:
    type:
    - boolean
    - 'null'
    default: true
    inputBinding:
      prefix: --correction
outputs:
  out_fastq1:
    type: File
    outputBinding:
      glob:
      - $(inputs.fastq1.nameroot).fastp.fastq
  out_fastq2:
    type:
    - File
    - 'null'
    outputBinding:
      glob:
      - $(inputs.fastq2.nameroot).fastp.fastq
  html_report:
    type: File
    outputBinding:
      glob:
      - fastp.html
  json_report:
    type: File
    outputBinding:
      glob:
      - fastp.json
arguments:
- prefix: -o
  valueFrom: $(inputs.fastq1.nameroot).fastp.fastq
- |
  ${
    if (inputs.fastq2){
      return '-O';
    } else {
      return '';
    }
  }
- |
  ${
    if (inputs.fastq2){
      return inputs.fastq2.nameroot + ".fastp.fastq";
    } else {
      return '';
    }
  }
baseCommand:
- fastp
```

and

```yaml
#!/usr/bin/env cwl-runner

cwlVersion: cwl:v1.0
class: CommandLineTool
label: MultiQC
doc: MultiQC is a tool to create a single report with interactive plots for multiple
  bioinformatics analyses across many samples.
inputs:
  files:
    type:
    - type: array
      items: File
    - 'null'
    doc: Files containing the result of quality analysis.
    inputBinding:
      position: 2
  directories:
    type:
    - type: array
      items: Directory
    - 'null'
    doc: Directories containing the result of quality analysis.
    inputBinding:
      position: 3
  report_name:
    type: string
    doc: Name of output report, without path but with full file name (e.g. report.html).
    default: multiqc_report.html
    inputBinding:
      position: 1
      prefix: -n
outputs:
  report:
    type: File
    outputBinding:
      glob:
      - '*.html'
baseCommand:
- multiqc
```

## Pipeline

Once the tools are created, we will create the pipeline itself using these two tools at **Projects > your\_project > Flow > Pipelines > CWL > Graphical**:

* On the Definition tab, go to the tool repository and drag and drop the two tools which you just created on the pipeline editor.
* Connect the JSON output of fastp to multiqc input by hovering over the middle of the round, blue connector of the output until the icon changes to a hand and then drag the connection to the first input of multiqc. You can use the magnification symbols to make it easier to connect these tools.
* Above the diagram, drag and drop two input FASTQ files and an output HTML file on to the pipeline editor and connect the blue markers to match the diagram below.

![fastp\_multiqc](/files/8a0HqOt3SvLwGqLfnvS0)

Relevant aspects of the pipeline:

* Both inputs are multivalue (as can be seen on the screenshot)
* Ensure that the step *fastp* has scattering configured: it scatters on both inputs using the scatter method 'dotproduct'. This means that as many instances of this step will be executed as there are pairs of FASTQ files. To indicate that this step is executed multiple times, the icons of both inputs have doubled borders.

### Important remark

Both input arrays (Read1 and Read2) must be matched. Currently an automatic sorting of input arrays is not supported. You have to take care of matching the input arrays which can be done in either one of two ways (besides the manual specification in the GUI):

* Invoke this pipeline in CLI using Bash functionality to sort the arrays
* Add a tool to the pipeline which will intake array of all FASTQ files, spread them on R1 and R2 suffixes, and sort them.

We will describe the second way in more detail. The tool will be based on public python Docker `docker.io/python:3.10` and have the following definition. In this tool we are providing the Python script *spread\_script.py* via Dirent [feature](https://help.ica.illumina.com/home/h-toolrepository#creating-your-first-tool-tips-and-tricks).

```yaml
#!/usr/bin/env cwl-runner

cwlVersion: v1.0
class: CommandLineTool
requirements:
- class: InlineJavascriptRequirement
- class: InitialWorkDirRequirement
  listing:
  - entry: "import argparse\nimport os\nimport json\n\n# Create argument parser\n\
      parser = argparse.ArgumentParser()\nparser.add_argument(\"-i\", \"--inputFiles\"\
      , type=str, required=True, help=\"Input files\")\n\n# Parse the arguments\n\
      args = parser.parse_args()\n\n# Split the inputFiles string into a list of file\
      \ paths\ninput_files = args.inputFiles.split(',')\n\n# Sort the input files\
      \ by the base filename\ninput_files = sorted(input_files, key=lambda x: os.path.basename(x))\n\
      \n\n# Separate the files into left and right arrays, preserving the order\n\
      left_files = [file for file in input_files if '_R1_' in os.path.basename(file)]\n\
      right_files = [file for file in input_files if '_R2_' in os.path.basename(file)]\n\
      \n# Print the left files for debugging\nprint(\"Left files:\", left_files)\n\
      \n# Print the left files for debugging\nprint(\"Right files:\", right_files)\n\
      \n# Ensure left and right files are matched\nassert len(left_files) == len(right_files),\
      \ \"Mismatch in number of left and right files\"\n\n    \n# Write the left files\
      \ to a JSON file\nwith open('left_files.json', 'w') as outfile:\n    left_files_objects\
      \ = [{\"class\": \"File\", \"path\": file} for file in left_files]\n    json.dump(left_files_objects,\
      \ outfile)\n\n# Write the right files to a JSON file\nwith open('right_files.json',\
      \ 'w') as outfile:\n    right_files_objects = [{\"class\": \"File\", \"path\"\
      : file} for file in right_files]\n    json.dump(right_files_objects, outfile)\n\
      \n"
    entryname: spread_script.py
    writable: false
label: spread_items
inputs:
  inputFiles:
    type:
      type: array
      items: File
    inputBinding:
      separate: false
      prefix: -i
      itemSeparator: ','
outputs:
  leftFiles:
    type:
      type: array
      items: File
    outputBinding:
      glob:
      - left_files.json
      loadContents: true
      outputEval: $(JSON.parse(self[0].contents))
  rightFiles:
    type:
      type: array
      items: File
    outputBinding:
      glob:
      - right_files.json
      loadContents: true
      outputEval: $(JSON.parse(self[0].contents))
baseCommand:
- python3
- spread_script.py
```

Now this tool can added to the pipeline before fastp step.


# Base Basics

Base is a genomics data aggregation and knowledge management solution suite. It is a secure and scalable integrated genomics data analysis solution which provides information management and knowledge mining. Refer to the [Base documentation](https://help.ica.illumina.com/project/p-base) for more details.

This tutorial provides an example for exercising the basic operations used with Base, including how to create a table, load the table with data, and query the table.

## Prerequisites

* A Platform Core project with access to Base
  * If you don't already have a project, please follow the instructions in the [Project documentation](/home/h-projects) to create a project.
* File to import
  * A tab delimited gene expression file ([sampleX.final.count.tsv](https://stratus-documentation-us-east-1-public.s3.amazonaws.com/DemoFiles/inputTSVfile.tsv)). Example format:

    ```tsv
    HES4-NM_021170-T00001  1392
    ISG15-NM_005101-T00002	46
    SLC2A5-NM_003039-T00003	14
    H6PD-NM_004285-T00004	30
    PIK3CD-NM_005026-T00005	200
    MTOR-NM_004958-T00006	156
    FBXO6-NM_018438-T00007	10
    MTHFR-NM_005957-T00008	154
    FHAD1-NM_052929-T00009	10
    PADI2-NM_007365-T00010	12
    ```

## Create table

Tables are components of databases that store data in a 2-dimensional format of columns and rows. Each row represents a new data record in the table; each column represents a field in the record. On Platform Core, you can use Base to create custom tables to fit your data. A schema definition defines the fields in a table. On Platform Core you can create a schema definition from scratch, or from a template. In this activity, you will create a table for RNAseq count data, by creating a schema definition from scratch.

1. Go to the **Projects > your\_project > Base > Tables** and enable Base by clicking on the **Enable** button.
2. Select **Add Table > New Table**.
3. Create your table

   1. To create your table from scratch, select Empty Table from the Create table from dropdown.
   2. Name your table *FeatureCounts*
   3. Uncheck the box next to *Include reference*, to exclude reference data from your table.
   4. Check the box next to *Edit as text*. This will reveal a text box that can be used to create your schema.
   5. Copy the schema text below and paste it in into the text box to create your schema.

   ```json
   {
     "Fields": [
       {
         "NAME_PATTERN": "[a-zA-Z][a-zA-Z0-9_]*",
         "Name": "TranscriptID",
         "Type": "STRING",
         "Mode": "REQUIRED",
         "Description": null,
         "DataResolver": null,
         "SubBluebaseFields": []
       },
       {
         "NAME_PATTERN": "[a-zA-Z][a-zA-Z0-9_]*",
         "Name": "ExpressionCount",
         "Type": "INTEGER",
         "Mode": "REQUIRED",
         "Description": null,
         "DataResolver": null,
         "SubBluebaseFields": []
       }
     ]
   }
   ```
4. Click the Save button

![save-table](/files/PszOwXLGJVbJNQSvNNBx)

## Upload data to load into your table

1. Upload **sampleX.final.count.tsv** file with the final count.
   1. Select **Data** tab (1) from the left menu.
   2. Click on the grey box (2) to choose the file to upload or drag and drop the **sampleX.final.count.tsv** into the grey box
   3. Refresh the screen (3)
   4. The uploaded file (4) will appear on the data page after successful upload.

![upload\_data](/files/tXJ4Hg1vpB8p7KMh3MYN)

## Create a schedule to load data into your table

Data can be loaded into tables manually or automatically. To load data automatically, you can set up a schedule. The schedule specifies which files’ data should be automatically loaded into a table, when those files are uploaded to Platform Core or created by an analyses on Platform Core. Active schedules will check for new files every 24 hours.

In this exercise, you will create a schedule to automatically load RNA transcript counts from **.final.count.tsv** files into the table you created above.

1. Go to **Projects > your\_project > Base > Schedule** and click the + Add New button.

![schedule-tab](/files/6roxPew1qdEnidnZgZp3)

2. Select the option to load the contents from files into a table.

<img src="/files/1P2OwZlax3YGDVl1XyHU" alt="schedule_load_content_from_file" width="206">

3. Create your schedule.
   1. Name your schedule **LoadFeatureCounts**
   2. Choose **Project** as the source of data for your table.
   3. To specify that data from **.final.count.tsv** files should be loaded into your table, enter **.final.count.tsv** in the **Search for a part of a specific ‘Orignal Name’ or Tag** text box.
   4. Specify your table as the one to load data into, by selecting your table (**FeatureCounts**) from the dropdown under **Target Base Table**.
   5. Under **Write preference**, select **Append to table**. New data will be appended to your table, rather than overwriting existing data in your table.
   6. The format of the **.final.count.tsv** files that will be loaded into your table are TSV/tab-delimited, and do not contain a header row. For the **Data format, Delimiter, and Header rows to skip** fields, use these values:
      * Data format: **TSV**
      * Delimiter: **Tab**
      * Header rows to skip: **0**
   7. Click the **Save** button

![create\_schedule](/files/aZvR3E3qrl4jwygrbp9G)

4. Highlight your schedule. Click the **Run** button to run your schedule now.

![run\_schedule](/files/sCsblQ8GUhWbRwgknRg3)

* It will take a short time to prepare and load data into your table.
  1. Check the status of your job on your **Projects > your\_project > Activity** page.
  2. Click the **BASE JOBS** tab to view the status of scheduled Base jobs.
  3. Click **BASE ACTIVITY** to view Base activity.

![activity](/files/pRep6BNP4VBFYZe3pGtV)

5. Check the data in the table.
   1. Go back to your **Projects > your\_project > Base > Tables** page.
   2. Double-click your table to view its details.
   3. You will land on the **SCHEMA DEFINITION** page.
   4. Click the PREVIEW tab to view the records that were loaded into your table.
   5. Click the DATA tab, to view a list of the files whose data has been loaded into your table.

![tablde\_preview\_data](/files/FjnMqczfsuqgESgNgmMr)

## Query a table

To request data or information from a Base table, you can run an SQL query. You can create and run new queries or saved queries.

In this activity, we will create and run a new SQL query to find out how many records (RNA transcripts) in your table have counts greater than 100.

1. Go to your **Projects > your\_project > Base > Query** page.

![query\_page](/files/q8aYyHsVpRj7tYUAkVK6)

```sql
SELECT TranscriptID,ExpressionCount FROM FeatureCounts WHERE ExpressionCount > 100;
```

2. Paste the above query into the **NEW QUERY** text box
3. Click the **Run Query** button to run your query
4. View your query results.
5. Save your query for future use by clicking the **Save Query** button. You will be asked to "Name" the query before clicking on the "Create" button.

![query\_results](/files/hfWKqyWVviKXZUKWjlkN)

## Export table data

Find the table you want to export on the "Tables" page under BASE. Go to the table details page by clicking twice on the table you want to export.

![table\_view](/files/FjnMqczfsuqgESgNgmMr)

Click on the **Export As File** icon and complete the required fields

1. **Name**: Name of the exported file
2. **Data Format**: A table can be exported in CSV and JSON format. The exported files can be compressed using GZIP, BZ2, DEFLATE or RAW\_DEFLATE.
   * CSV Format: In addition to Comma, the file can be Tab, Pipe or Custom character delimited.
   * JSON Format: Selecting JSON format exports the table in a text file containing a JSON object for each entry in the table. This is the standard snowflake behavior.

![json-format](/files/iMG0O4Xv2nKtb6Q30blP)

3. **Export to single/multiple files**: This option allows the export of a table as a single (large) file or multiple (smaller) files. If "Export to multiple files" is selected, a user can provide "Maximum file size (in bytes)" for exported files. The default value is 16,000,000 bytes but can be increased to accommodate larger files. The maximum file size supported is 5 GB.


# Base: SnowSQL

You can access the databases and tables within the Base module using snowSQL command-line interface. This is useful for external collaborators who do not have access to Platform Core functionalities. In this tutorial we will describe how to obtain the token and use it for accessing the Base module. This tutorial does not cover how to install and configure snowSQL.

## Obtaining OAuth token and URL

**Once the Base module has been enabled within a project**, the following details are shown in **Projects > your\_project > Project Settings > Details**.

![base-enabled-oauth](/files/Dxlgo8ISZ0lv6behFdNZ)

After clicking the button `Create OAuth access token`, the pop-up authenticator is displayed.

![base-oauth-token](/files/k0lSlWAqfRiFv2LkHa7V)

After clicking the button `Generate snowSQL command` the pop-up authenticator presents the snowSQL command.

![base-oauth-command](/files/c73JSPQTogZ1iu1SdN8a)

Copy the snowSQL command and run it in the console to log in.

You can also get the OAuth access token via API by providing \<PROJECT ID> and \<YOUR KEY>.

### Example:

API Call:

```
curl -X 'POST' \
  'https://ica.illumina.com/ica/rest/api/projects/<PROJECT ID>/base:connectionDetails' \
  -H 'accept: application/vnd.illumina.v3+json' \
  -H 'X-API-Key: <YOUR KEY>' 
```

Response

```
{
  "authenticator": "oauth",
  "accessToken": "XXXXXXXXXX",
  "dnsName": "use1sf01.us-east-1.snowflakecomputing.com",
  "userPrincipalName": "xxxxx",
  "databaseName": "xxxxx",
  "schemaName": "xxx",
  "warehouseName": "xxxxxx",
  "roleName": "xxx"
}
```

Template snowSQL:

```
snowsql -a use1sf01.us-east-1 -u <userPrincipalName> --authenticator=oauth -r <roleName> -d <databaseName> -s PUBLIC -w <warehouseName> --token="<accessToken>"
```

Now you can perform a variety of tasks such as:

1. Querying Data: execute SQL queries against tables, views, and other database objects to retrieve data from the Snowflake data warehouse.
2. Creating and Managing Database Objects: create tables, views, stored procedures, functions, and other database objects in Snowflake. you can also modify and delete these objects as needed.
3. Loading Data: load data into Snowflake from various sources such as local files, AWS S3, Azure Blob Storage, or Google Cloud Storage.

Overall, snowSQL CLI provides a powerful and flexible interface to work with Snowflake, allowing external users to manage data warehouse and perform a variety of tasks efficiently and effectively without access to Platform Core.

### Example Queries:

Show all tables in the database:

```
>SHOW TABLES;
```

Create a new table:

```
create TABLE demo1(sample_name VARCHAR, count INT);
```

List records in a table:

```
SELECT * FROM demo1;
```

Load data from a file: To load data from a file, you can start by create a staging area in the internal storage using the following commend:

```
>CREATE STAGE myStage;
```

You can then upload the local file to the internal storage using the following command:

```
> PUT file:///path/to/data.tsv @myStage;
```

You can check if the file was uploaded properly using LIST command:

```
> LIST @myStage;
```

Finally, Load data by using COPY TO command. The command assumes the data.tsv is a tab delimited file. You can easily modify the following command to import JSON file setting TYPE=JSON.

```
> COPY INTO demo1(sample_name, count) FROM @mystage/data.tsv FILE_FORMAT = (TYPE = 'CSV' FIELD_DELIMITER = '\t');
```

Load data from a string: If you have data as JSON string, you can import the data into the tables using following commands.

```
> SET myJSON_str = '{"sample_name": "from-json-str", "count": 1}';
> INSERT INTO demo1(sample_name, count)
> SELECT
    PARSE_JSON($myJSON_str):sample_name::STRING,
    PARSE_JSON($myJSON_str):count::INT
```

Load data into specific columns: If you want to load sample\_name into the table, you can remove the "count" from the column and the value list as below:

```
> SET myJSON_str = '{"sample_name": "from-json-str", "count": 1}';
> INSERT INTO demo1(sample_name)
  SELECT
    PARSE_JSON($myJSON_str):sample_name::STRING;
```

List the views of the database to which you are connected. As shared database and catalogue views are created within the project database, they will be listed. However, it does not show views which are granted via another database, role or from bundles.

```
>SHOW VIEW;
```

Show grants, both directly on the tables and views and grants to roles which in turn have grants on tables and views.

```
>SHOW GRANTS;
```


# Base: Access Tables via Python

You can access the databases and tables within the Base module using Python from your local machine. Once retrieved as e.g. *pandas* object, the data can be processed further. In this tutorial, we will describe how you could create a Python script which will retrieve the data and visualize it using *Dash* framework. The script will contain the following parts:

* Importing dependencies and variables.
* Function to fetch the data from Base table.
* Creating and running the Dash app.

## Importing dependencies and variables

This part of the code imports the dependencies which have to be installed on your machine (possibly with *pip*). Furthermore, it imports the variables API\_KEY and PROJECT\_ID from the file named *config*.

```python
from dash import Dash, html, dcc, callback, Output, Input
import plotly.express as px

from config import API_KEY, PROJECT_ID
import requests
import snowflake.connector
import pandas as pd
```

## Function to fetch the data from Base table

We will be creating a function called *fetch\_data* to obtain the data from Base table. It can be broken into several logically separated parts:

* Retrieving the token to access the Base table together with other variables using API.
* Establishing the connection using the token.
* SQL query itself. In this particular example, we are extracting values from two tables *Demo\_Ingesting\_Metrics* and *BB\_PROJECT\_PIPELINE\_EXECUTIONS\_DETAIL*. The table *Demo\_Ingesting\_Metrics* contains various metrics from DRAGEN analyses (e.g. the number of bases with quality at least 30 *Q30\_BASES*) and metadata in the column *ica* which needs to be flattened to access the value *Execution\_reference*. Both tables are joined on this *Execution\_reference* value.
* Fetching the data using the connection and the SQL query.

Here is the corresponding snippet:

```python
def fetch_data():
    # Your data fetching and processing code here
    # retrieving the Base oauth token
    url = 'https://ica.illumina.com/ica/rest/api/projects/' + PROJECT_ID +  '/base:connectionDetails'

    # set the API headers
    headers = {
                'X-API-Key': API_KEY,
                'accept': 'application/vnd.illumina.v3+json'
                }

    response = requests.post(url, headers=headers)
    ctx = snowflake.connector.connect(
        account=response.json()['dnsName'].split('.snowflakecomputing.com')[0],
        authenticator='oauth',
        token=response.json()['accessToken'], 
        database=response.json()['databaseName'],
        role=response.json()['roleName'],
        warehouse=response.json()['warehouseName']
    )
    cur = ctx.cursor()
    sql = '''
    WITH flattened_Demo_Ingesting_Metrics AS (
        SELECT 
            flattened.value::STRING AS execution_reference_Demo_Ingesting_Metrics,
            t1.SAMPLEID,
            t1.VARIANTS_TOTAL_PASS,
            t1.VARIANTS_SNPS_PASS,
            t1.Q30_BASES,
            t1.READS_WITH_MAPQ_3040_PCT
        FROM 
            Demo_Ingesting_Metrics t1,
            LATERAL FLATTEN(input => t1.ica) AS flattened
        WHERE 
            flattened.key = 'Execution_reference'
    ) SELECT 
        f.execution_reference_Demo_Ingesting_Metrics,
        f.SAMPLEID,
        f.VARIANTS_TOTAL_PASS,
        f.VARIANTS_SNPS_PASS,
        t2."EXECUTION_REFERENCE",
        t2.END_DATE,
        f.Q30_BASES,
        f.READS_WITH_MAPQ_3040_PCT
    FROM 
        flattened_Demo_Ingesting_Metrics f
    JOIN 
        BB_PROJECT_PIPELINE_EXECUTIONS_DETAIL t2
    ON 
        f.execution_reference_Demo_Ingesting_Metrics = t2."EXECUTION_REFERENCE";
    '''

    cur.execute(sql)
    data = cur.fetch_pandas_all()
    return data

df = fetch_data()
```

## Creating and running the Dash app

Once the data is fetched, it is visualized in an app. In this particular example, a scatter plot is presented with END\_DATE as **x** axis and the choice of the customer from the dropdown as **y** axis.

```python

app = Dash(__name__)
#server = app.server


app.layout = html.Div([
    html.H1("My Dash Dashboard"),
    
    html.Div([
        html.Label("Select X-axis:"),
        dcc.Dropdown(
            id='x-axis-dropdown',
            options=[{'label': col, 'value': col} for col in df.columns],
            value=df.columns[5]  # default value
        ),
        html.Label("Select Y-axis:"),
        dcc.Dropdown(
            id='y-axis-dropdown',
            options=[{'label': col, 'value': col} for col in df.columns],
            value=df.columns[2]  # default value
        ),
    ]),
    
    dcc.Graph(id='scatterplot')
])


@callback(
    Output('scatterplot', 'figure'),
    Input('y-axis-dropdown', 'value')
)
def update_graph(value):
    return px.scatter(df, x='END_DATE', y=value, hover_name='SAMPLEID')

if __name__ == '__main__':
    app.run(debug=True)
```

Now we can create a single Python script called *dashboard.py* by concatenating the snippets and running it. The dashboard will be accessible in the browser on your machine.


# Bench ICA Python Library

This tutorial demonstrates how to use the Platform Core Python library packaged with the JupyterLab image for Bench Workspaces.

See the [JupyterLab documentation](/project/p-bench/bench-jupyterlab) for details about the JupyterLab docker image provided by Illumina.

The tutorial will show how authentication to the Platform Core API works and how to search, upload, download and delete data from a project into a Bench Workspace. The python code snippets are written for compatibility with a Jupyter Notebook.

## Python modules

Navigate to **Bench > Workspaces** and click **Enable** to enable workspaces. Select **+New Workspace** to create a new workspace. Fill in the required details and select JupyterLab for the Docker image. Click **Save and Start** to open the workspace. The following snippets of code can be pasted into the workspace you've created.

This snippet defines the required python modules for this tutorial:

```python
# Wrapper modules
import icav2
from icav2.api import project_data_api
from icav2.model.problem import Problem
from icav2.model.project_data import ProjectData

# Helper modules
import random
import os
import requests
import string
import hashlib
import getpass
```

## Authentication

This snippet shows how to authenticate using the following methods:

* Platform Core username and password
* Platform Core API token

```python
# Authenticate using User credentials
username = input("Core Username")
password = getpass.getpass("Core Password")
tenant = input("Core Tenant name")
url = os.environ['ICA_URL'] + '/rest/api/tokens'
r = requests.post(url, data={}, auth=(username,password),params={'tenant':tenant})
token = None
apiClient = None
if r.status_code == 200:
    token = r.content
    configuration = icav2.Configuration(
        host = os.environ['ICA_URL'] + '/rest',
        access_token = str(r.json()["token"])
        )
    apiClient = icav2.ApiClient(configuration, header_name="Content-Type",header_value="application/vnd.illumina.v3+json")
    print("Authenticated to %s" % str(os.environ['ICA_URL']))
else:
    print("Error authenticating to %s" % str(os.environ['ICA_URL']))
    print("Response: %s" % str(r.status_code))

## Authenticate using ICA API TOKEN
configuration = icav2.Configuration(
    host = os.environ['ICA_URL'] + '/rest'
)
configuration.api_key['ApiKeyAuth'] = getpass.getpass()
apiClient = icav2.ApiClient(configuration, header_name="Content-Type",header_value="application/vnd.illumina.v3+json")
```

## Data Operations

These snippets show how to manage data in a project. Operations shown are:

* Create a Project Data API client instance
* List all data in a project
* Create a data element in a project
* Upload a file to a data element in a project
* Download a data element from a project
* Search for matching data elements in a project
* Delete matching data elements in a project

```python
# Retrieve project ID from the Bench workspace environment
projectId = os.environ['ICA_PROJECT']
```

```python
# Create a Project Data API client instance
projectDataApiInstance = project_data_api.ProjectDataApi(apiClient)
```

### List Data

```python
# List all data in a project
pageOffset = 0
pageSize = 30
try:
    projectDataPagedList = projectDataApiInstance.get_project_data_list(project_id = projectId, page_size = str(pageSize), page_offset = str(pageOffset))
    totalRecords = projectDataPagedList.total_item_count
    while pageOffset*pageSize < totalRecords:
        for projectData in projectDataPagedList.items:
            print("Path: "+projectData.data.details.path + " - Type: "+projectData.data.details.data_type)
        pageOffset = pageOffset + 1
except icav2.ApiException as e:
    print("Exception when calling ProjectDataAPIApi->get_project_data_list: %s\n" % e)
```

### Create Data

```python
# Create data element in a project
data = icav2.model.create_data.CreateData(name="test.txt",data_type = "FILE")

try:
    projectData = projectDataApiInstance.create_data_in_project(projectId, create_data=data)
    fileId = projectData.data.id
except icav2.ApiException as e:
    print("Exception when calling ProjectDataAPIApi->create_data_in_project: %s\n" % e)
```

### Upload Data

```python
## Upload a local file to a data element in a project
# Create a local file in a Bench workspace
filename = '/tmp/'+''.join(random.choice(string.ascii_lowercase) for i in range(10))+".txt"
content = ''.join(random.choice(string.ascii_lowercase) for i in range(100))
f = open(filename, "a")
f.write(content)
f.close()

# Calculate MD5 hash (optional)
localFileHash = md5Hash = hashlib.md5((open(filename, 'rb').read())).hexdigest()

try:
    # Get Upload URL
    upload = projectDataApiInstance.create_upload_url_for_data(project_id = projectId, data_id = fileId)
    # Upload dummy file
    files = {'file': open(filename, 'r')}
    data = open(filename, 'r').read()
    r = requests.put(upload.url , data=data)
except icav2.ApiException as e:
    print("Exception when calling ProjectDataAPIApi->create_upload_url_for_data: %s\n" % e)

# Delete local dummy file
os.remove(filename)
```

### Download Data

```python
## Download a data element from a project
try:
    # Get Download URL 
    download = projectDataApiInstance.create_download_url_for_data(project_id=projectId, data_id=fileId)

    # Download file
    filename = '/tmp/'+''.join(random.choice(string.ascii_lowercase) for i in range(10))+".txt"
    r = requests.get(download.url)
    open(filename, 'wb').write(r.content)

    # Verify md5 hash
    remoteFileHash = hashlib.md5((open(filename, 'rb').read())).hexdigest()
    if localFileHash != remoteFileHash:
        print("Error: MD5 mismatch")

    # Delete local dummy file
    os.remove(filename)
except icav2.ApiException as e:
    print("Exception when calling ProjectDataAPIApi->create_download_url_for_data: %s\n" % e)
```

### Search for Data

```python
# Search for matching data elements in a project
try:
    projectDataPagedList = projectDataApiInstance.get_project_data_list(project_id = projectId, full_text="test.txt")
    for projectData in projectDataPagedList.items:
        print("Path: " + projectData.data.details.path + " - Name: "+projectData.data.id + " - Type: "+projectData.data.details.data_type)
except icav2.ApiException as e:
    print("Exception when calling ProjectDataAPIApi->get_project_data_list: %s\n" % e)
```

### Delete Data

```python
# Delete matching data elements in a project
try:
    projectDataPagedList = projectDataApiInstance.get_project_data_list(project_id = projectId, full_text="test.txt")
    for projectData in projectDataPagedList.items:
        print("Deleting file "+projectData.data.details.path)  
        projectDataApiInstance.delete_data(project_id = projectId, data_id = projectData.data.id)
except icav2.ApiException as e:
    print("Exception %s\n" % e)
```

## Base Operations

These snippets show how to get a connection to a base database and run an example query. Operations shown are:

* Create a python jdbc connection
* Create a table
* Insert data into a table
* Query the table
* Delete the table

Snowflake Python API documentation can be found [here](https://docs.snowflake.com/en/user-guide/python-connector.html)

This snipppet defines the required python modules for this tutorial:

```python
# API modules
import icav2
from icav2.api import project_base_api
from icav2.model.problem import Problem
from icav2.model.base_connection import BaseConnection

# Helper modules
import os
import requests
import getpass
import snowflake.connector
```

```python
# Retrieve project ID from the Bench workspace environment
projectId = os.environ['ICA_PROJECT']
```

```python
# Create a Project Base API client instance
projectBaseApiInstance = project_base_api.ProjectBaseApi(apiClient)
```

### Get Base Access Credentials

```python
# Get a Base Access Token
try:
    baseConnection = projectBaseApiInstance.create_base_connection_details(project_id = projectId)
except icav2.ApiException as e:
    print("Exception when calling ProjectBaseAPIApi->create_base_connection_details: %s\n" % e)
## Create a python jdbc connection
ctx = snowflake.connector.connect(
    account=os.environ["ICA_SNOWFLAKE_ACCOUNT"],
    authenticator=baseConnection.authenticator,
    token=baseConnection.access_token, 
    database=os.environ["ICA_SNOWFLAKE_DATABASE"],
    role=baseConnection.role_name,
    warehouse=baseConnection.warehouse_name
)
ctx.cursor().execute("USE "+os.environ["ICA_SNOWFLAKE_DATABASE"])
```

### Create a Table

```python
## Create a Table
tableName = "test_table"
ctx.cursor().execute("CREATE OR REPLACE TABLE " + tableName + "(col1 integer, col2 string)")
```

### Add Table Record

```python
## Insert data into a table
ctx.cursor().execute(
        "INSERT INTO " + tableName + "(col1, col2) VALUES " + 
        "    (123, 'test string1'), " + 
        "    (456, 'test string2')")
```

### Query Table

```python
## Query the table
cur = ctx.cursor()
try:
    cur.execute("SELECT * FROM "+tableName)
    for (col1, col2) in cur:
        print('{0}, {1}'.format(col1, col2))
finally:
    cur.close()
```

### Delete Table

```python
# Delete the table
ctx.cursor().execute("DROP TABLE " + tableName);
```


# API Beginner Guide

## API Basics

Any operation from the Platform Core graphical user interface can also be performed with the API.

The following are some basic examples on how to use the API. These examples are based on using Python as programming language. For other languages, please see their native documentation on API usage.

### Prerequisites

* An installed copy of Python. (<https://www.python.org/>)
* The package installer for python (pip) (<https://pip.pypa.io/>)
* Having the python requests library installed (`pip install requests`)

### Authenticating

One of the easiest authentication methods is by means of API keys. To generate an API key, refer to the [Get Started](broken://spaces/7GiJwg33pKa8eXmwcle7/pages/0arMqMwXm4yiUFrZ90Qo) section. This key is then used in your Python code to authenticate the API calls. It is best practice to regularly update your API keys.

API keys are valid for a single user, so any information you request is for that user to which the key belongs. For this reason, it is best practice to create a dedicated API user so you can manage the access rights for the API by managing those user rights.

### API Reference

There is a dedicated [API Reference](https://ica.illumina.com/ica/api/swagger/index.html) where you can enter your API key and try out the different Platform Core API commands and get an overview of the available parameters.

### Converting curl to Python

The examples on the [API Reference](https://ica.illumina.com/ica/api/swagger/index.html) page use curl (Client URL) while Python uses Python requests. There are a number of online tools to automatically convert from curl to python.

To get the curl command,

1. Look up the endpoint you want to use on the API reference page.
2. Select `Try it out`.
3. Enter the necessary parameters.
4. Select `Execute`.
5. Copy the resulting curl command.

> **Never paste your API authentication key into online tools when performing curl conversion as this poses a significant security risk.**

In the most basic form, the curl command

`curl my.curlcommand.com`

becomes

```
import requests
response = requests.get('http://my.curlcommand.com')
```

You will see the following options in the curl commands on the [API Reference](https://ica.illumina.com/ica/api/swagger/index.html) page.

`-H` means header.

`-X` means the string is passed "as is" without interpretation.

```
curl -X 'GET' 'https://my.curlcommand.com' -H 'HeaderName: HeaderValue'
```

becomes

```
import requests
headers = {
    'HeaderName': 'HeaderValue',
}
response = requests.get('https://my.curlcommand.com', headers=headers)
```

## Simple API Examples

### **Request a list of event codes**

This is a straightforward request without parameters which can be used to to verify your connection.

The API call is

`response = requests.get('https://ica.illumina.com/ica/rest/api/eventcodes', headers={'X-API-Key': '<your_generated_API_key>'})`

In this example, the API key is part of your API call, which means you must update all API calls when the key changes. A better practice is to put this API key in the headers so it is easier to maintain. The full code then becomes

```
# The requests library will allow you to make HTTP requests.
import requests

# Replace <your_generated_API_key> with your actual generated API key here.
headers = {
    'X-API-Key': '<your_generated_API_key>',
}

# Store the API request in response.
response = requests.get('https://ica.illumina.com/ica/rest/api/eventCodes', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Response status code: ", response.status_code)

# Display the data from the request.
print(response.json())
```

### **Pretty-printing the result**

The list of application codes was returned as a single line, which makes it difficult to read, so let's pretty-print the result.

```
# The requests library will allow you to make HTTP requests.
import requests

# JSON will allow us to format and interpret the output.
import json

# Replace <your_generated_API_key> with your actual generated API key here.
headers = {
    'X-API-Key': '<your_generated_API_key>',
}

# Store the API request in response.
response = requests.get('https://ica.illumina.com/ica/rest/api/eventCodes', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Print JSON data in readable format with indentation and sorting.
print(json.dumps(My_API_Data, indent=3, sort_keys=True))
```

### **Retrieving a list of projects**

Now that we are able to retrieve information with the API, we can use it for a more practical request like retrieving a list of projects. This API request can also take parameters.

#### Retrieve all projects

First, we pass the request without parameters to retrieve all projects.

```
# The requests library will allow you to make HTTP requests.
import requests

# JSON will allow us to format and interpret the output.
import json

# Replace <your_generated_API_key> with your actual generated API key here.
headers = {
    'X-API-Key': '<your_generated_API_key>',
}

# Store the API request in response.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Print JSON data in readable format with indentation and sorting.
print(json.dumps(My_API_Data, indent=3, sort_keys=True))
```

#### Single parameter

The easiest way to pass a parameter is by appending it to the API request. The following API request will list the projects with a filter on CAT as user tag.

`response = requests.get('https://ica.illumina.com/ica/rest/api/projects?userTags=CAT', headers=headers)`

#### Multiple parameters

If you only want entries that have both the tags CAT and WOLF, you would append them like this:

`response = requests.get('https://ica.illumina.com/ica/rest/api/projects?userTags=CAT&userTags=WOLF', headers=headers)`

### **Copying Data**

To copy data, you need to know:

* Your generated API key.
* The dataId of the files and folders which you want to copy (their syntax is fil.hexadecimal\_identifier and fol.hexadecimal\_identifier). You can select a file or folder in the GUI and select it to see the Id (Projects > your\_project > Data > your\_file > Data details > Id) or you can use the `/api/projects/{projectId}/data` endpoint.
* The destination project to which you want to copy the data.
* The destination folder within the destination project to which you want to copy the data (fol.hexadecimal\_identifier).
* What to do when the destination files or folders already exist (OVERWRITE, SKIP or RENAME).

The full code will then be as follows:

```
# The requests library will allow you to make HTTP requests.
import requests

# Fill out your generated API key.
headers = {
    'accept': 'application/vnd.illumina.v3+json',
    'X-API-Key': '<your_generated_API_key>',
    'Content-Type': 'application/vnd.illumina.v3+json',
}

# Enter the files and folders, the destination folder, and the action to perform when the destination data already exists.
data = '{"items": [{"dataId": "fil.0123456789abcdef"}, {"dataId": "fil.735040537abcdef"}], "destinationFolderId": "fol.1234567890abcdef", "copyUserTags": true,"copyTechnicalTags": true,"copyInstrumentInfo": true,"actionOnExist": "SKIP"}'

# Replace <Project_Identifier> with the actual identifier of the destination project.
response = requests.post(
    'https://ica.illumina.com/ica/rest/api/projects/**<Project_Identifier>**/dataCopyBatch',
    headers=headers,
    data=data,
)

# Display the response status code.
print("Response status code: ", response.status_code) 
```

## Combined API Example - Running a Pipeline

Now that we have done individual API requests, we can combine them and use the output of one request as input for the next request. When you want to run a pipeline, you need a number of input parameters. In order to obtain these parameters, you need to make a number of API calls first and use the returned results as part of your request to run the pipeline. In the examples below, we will build up the requests one by one so you can run them individually first to see how they work. These examples only follow the happy path to keep them as simple as possible. If you program them for a full project, remember to add error handling. You can also use the GUI to get all the parameters or write them down after performing the individual API calls in this section. Then, you can build your final API call with those values fixed.

### Initialization

This block must be added at the beginning of your code

```
# The requests library will allow you to make HTTP requests.
import requests

# JSON will allow us to format and interpret the output.
import json

# Replace <your_generated_API_key> with your actual generated API key here.
headers = {
    'X-API-Key': '<your_generated_API_key>',
}
```

### Look for a project in the list of Projects

Previously, we already requested a list of all projects, now we add a search parameter to look for a project called **MyProject**. (Replace MyProject with the name of the project you want to look for).

```
# Store the API request in response. Here we look for a project called "MyProject".
response = requests.get('https://ica.illumina.com/ica/rest/api/projects?search=MyProject', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Print JSON data in readable format with indentation and sorting.
print(json.dumps(My_API_Data, indent=3, sort_keys=True))
```

Now that we have found our project by name, we need to get the unique project id, which we will use in the combined requests. To get the id, we add the following line to the end of the code above.

```
print(My_API_Data['items'][0]['id'])
```

Syntax \['items']\[0]\['id'] means we look for the items list, 0 means we take the first entry (as we presume our filter was accurate enough to only return the correct result and we don't have duplicate project names) and id means we take the data from the id field. Similarly, you can build other expressions to give you the data you want to see, such as \['items']\[0]\['urn'] to get the urn or \['items']\[0]\['tags']\['userTags'] to get the list of user tags.

Once we have the identifier we need, we add it to a variable which we will call Project\_Identifier in our examples.

```
# Get the project identifier.
Project_Identifier = My_API_Data['items'][0]['id']
```

### Retrieve the Pipelines of your Project

Once we have the identifier of our project, we can fill it out in the request to list the pipelines which are part of our project.

```
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/pipelines?', headers=headers)
```

This will give us all the available pipelines for that project. As we will only want to run a single pipeline, we can search for our pipeline, which in this example will be the basic\_pipeline. Unfortunately, this API call has no direct search parameter, so when we get the list of pipelines, we will look for the id and store that in a variable which we will call Pipeline\_Identifier in our examples as follows:

```
# Find Pipeline
# Store the API request in response. Here we look for the list of pipelines in MyProject.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/pipelines', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find Pipeline Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Store the list of pipelines for further processing.
pipelineslist = json.dumps(My_API_Data)

# Set "basic_pipeline" as the pipeline to search for. Replace this with your target pipeline.
target_pipeline = "basic_pipeline"
found_pipeline = None

# Look for the code to match basic_pipeline and store the ID.
for item in My_API_Data['items']:
    if 'pipeline' in item and item['pipeline'].get('code') == target_pipeline:
        found_pipeline = item['pipeline']
        Pipeline_Identifier = found_pipeline['id']
        break
print("Pipeline Identifier: " + Pipeline_Identifier)
```

### Find which parameters the Pipeline needs.

Once we know the project identifier and the pipeline identifier, we can create an API request to retrieve the list of input parameters which are needed for the pipeline. We will consider a simple pipeline which only needs a file as input. If your pipeline has more input parameters, you will need to set those as well.

```
# Find Parameters
# Store the API request in response. Here we look for the Parameters in basic_pipeline
response = requests.get('https://ica.illumina.com/ica/rest/api/pipelines/'+(Pipeline_Identifier)+'/inputParameters', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find Parameters Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Get the parameters and store in the Parameters variable.
Parameters = My_API_Data['items'][0]['code']
print("Parameters: ",Parameters)
```

### Find the Storage Size to use for the analysis.

Here we will look for the id of the extra small storage size. This is done with the 0 in the My\_API\_Data\['items']\[0]\['id']

```
# Store the API request in response. Here we look for the analysis storage size.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/analysisStorages', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find analysisStorages Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Get the storage size. We will select the smallest size.
Storage_Size = My_API_Data['items'][0]['id']
print("Storage_Size: ",Storage_Size)
```

### Find the files to use as input for your pipeline.

Now we will look for a file "testExample" which we want to use as input and store the file id.

```
# Get Input File
# Store the API request in response. Here we look for the Files testExample.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/data?fullText=testExample', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find input file Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data
My_API_Data = response.json()

# Get the first file ID.
InputFile = My_API_Data['items'][0]['data']['id']
print("InputFile id: ",InputFile)
```

### Start the Pipeline.

Finally, we can run the analysis with parameters filled out.

```
Postheaders = {
    'accept': 'application/vnd.illumina.v4+json',
    'X-API-Key': '<your_generated_API_key>',
    'Content-Type': 'application/vnd.illumina.v4+json',
}

data = '{"userReference":"api_example","pipelineId":"'+(Pipeline_Identifier)+'","analysisStorageId":"'+(Storage_Size)+'","analysisInput":{"inputs":[{"parameterCode":"'+(Parameters)+'","dataIds":["'+(InputFile)+'"]}]}}'

response = requests.post(
    'https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/analysis:nextflow',headers=Postheaders,data=data,
)
```

### Complete code example

```
# The requests library will allow you to make HTTP requests.
import requests

# JSON will allow us to format and interpret the output.
import json

# Replace <your_generated_API_key> with your actual generated API key here.
headers = {
    'X-API-Key': '<your_generated_API_key>',
}

# Replace <your_generated_API_key> with your actual generated API key here.
Postheaders = {
    'accept': 'application/vnd.illumina.v4+json',
    'X-API-Key': '<your_generated_API_key>',
    'Content-Type': 'application/vnd.illumina.v4+json',
}

# Find project
# Store the API request in response. Here we look for a project called "MyProject".
response = requests.get('https://ica.illumina.com/ica/rest/api/projects?search=MyProject', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find Project response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Get the project identifier.
Project_Identifier = My_API_Data['items'][0]['id']
print("Project_Identifier: ",Project_Identifier)

# Find Pipeline
# Store the API request in response. Here we look for the list of pipelines in MyProject.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/pipelines', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find Pipeline Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Store the list of pipelines for further processing.
pipelineslist = json.dumps(My_API_Data)

# Set "basic_pipeline" as the pipeline to search for. Replace this with your target pipeline.
target_pipeline = "basic_pipeline"
found_pipeline = None

# Look for the code to match basic_pipeline and store the ID.
for item in My_API_Data['items']:
    if 'pipeline' in item and item['pipeline'].get('code') == target_pipeline:
        found_pipeline = item['pipeline']
        Pipeline_Identifier = found_pipeline['id']
        break
print("Pipeline Identifier: " + Pipeline_Identifier)

# Find Parameters
# Store the API request in response. Here we look for the Parameters in basic_pipeline.
response = requests.get('https://ica.illumina.com/ica/rest/api/pipelines/'+(Pipeline_Identifier)+'/inputParameters', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find Parameters Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Get the parameters and store in the Parameters variable.
Parameters = My_API_Data['items'][0]['code']
print("Parameters: ",Parameters)

# Get Storage Size
# Store the API request in response. Here we look for the analysis storage size.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/analysisStorages', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find analysisStorages Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Get the storage size. We will select the smallest size.
Storage_Size = My_API_Data['items'][0]['id']
print("Storage_Size: ",Storage_Size)

# Get Input File
# Store the API request in response. Here we look for the Files testExample.
response = requests.get('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/data?fullText=testExample', headers=headers)

# Display the response status code. Code 200 means the request succeeded.
print("Find input file Response status code: ", response.status_code)

# Put the JSON data from the response in My_API_Data.
My_API_Data = response.json()

# Get the first file ID.
InputFile = My_API_Data['items'][0]['data']['id']
print("InputFile id: ",InputFile)

# Finally, we can run the analysis with parameters filled out.
data = '{"userReference":"api_example","pipelineId":"'+(Pipeline_Identifier)+'","tags":{"technicalTags":[],"userTags":[],"referenceTags":[]},"analysisStorageId":"'+(Storage_Size)+'","analysisInput":{"inputs":[{"parameterCode":"'+(Parameters)+'","dataIds":["'+(InputFile)+'"]}]}}'
print (data)
response = requests.post('https://ica.illumina.com/ica/rest/api/projects/'+(Project_Identifier)+'/analysis:nextflow',headers=Postheaders,data=data,)
print("Post Response status code: ", response.status_code)
```


# Launching a DRAGEN Pipeline

This guide covers launching, monitoring, and debugging DRAGEN pipelines using the DRAGEN Germline Whole Genome pipeline as example.

## Prerequisites

Before launching the pipeline, ensure you have the following in place:

* **Platform Core Project** — You must have an existing project in Platform Core. If you need to create a new project, follow the insctructions described in [Projects](/home/h-projects#create-new-project).
* **DRAGEN Bundle** — The DRAGEN bundle must be linked to your project. See [Linking Bundles](/home/h-bundles#linking-an-existing-bundle-to-a-project). This provides bundled references, pipelines, and demo data.
* **Input Data** — Upload your sequencing data (FASTQ, ORA, BAM, or CRAM files) to your [project](/project/p-data), or use the demo data provided in "Illumina DRAGEN Germline Demo Data."

## Launching via the Platform Core GUI

{% stepper %}
{% step %}

### Start a New Analysis

1. Navigate to **Projects > your\_project > Flow > Pipelines**.
2. Select **DRAGEN\_Germline\_Whole\_Genome**.
3. (Optional) Read the **Pipeline** **Documentation** page to find out more information about the pipeline, including its changelog and additional resources.
4. Click **Start analysis**.
5. Enter a **User Reference** (a meaningful name for this analysis run) and select a **Subscription** from the Pricing drop-down.
   {% endstep %}

{% step %}

### Configure Inputs

Select your **Input Type** (FASTQ GZ, FASTQ ORA, BAM, or CRAM) and provide your input files:

* **FASTQs / ORAs** — Select your sequencing files. Multiple samples may be provided. The pipeline automatically parses filenames to determine FASTQ pairs and sample groupings. The RGSM is taken from the filename up to `_SX` and the suffix after `_RX`. `_LXXX` denotes the lane number and `_RX` denotes the read number.

{% hint style="info" %}
To override the auto-detected sample names, you can provide a **FASTQ List** CSV containing the filenames with your own specified RGSM values.
{% endhint %}

* **BAMs / CRAMs** — Select your alignment files. Map/Align can be turned off if realignment is not desired.

Select a **Reference** genome from the drop-down. The default is **Homo sapiens \[1000 Genomes] hg38 v6 Pangenome**. Expand the drop-down for the full list of bundled references, or select **Custom** and provide your own reference hash table.
{% endstep %}

{% step %}

### Configure Analysis Options

The input form provides options that vary by pipeline. Common sections include Map/Align, Variant Calling, CNV, SV, Variant Annotation, and Advanced Settings. Some pipelines also expose sections such as UMI, HLA Typing, Methylation, Fingerprint Checking, Targeted Callers, or Beta Features. Each field includes built-in help text describing its purpose and valid values.

For a standard WGS germline run, the defaults are suitable for most use cases — see [Analysis Settings](#analysis-settings). Review and adjust any options as needed for your experiment.
{% endstep %}

{% step %}

### Launch

Review your settings and click **Start analysis**.
{% endstep %}
{% endstepper %}

## Launching the Pipeline via CLI

If you have the [CLI](/command-line-interface/cli-indexcommands) installed on your system, you can also use the commands below to work with pipelines. If you do not have an active CLI, please follow [these instructions](/command-line-interface/cli-installation) first.

{% hint style="info" %}
For any `icav2` CLI command, you can append `--help` to see a list of available optional settings.
{% endhint %}

### Accessing your Pipeline

1. List your projects with

```bash
icav2 projects list
```

2. Enter your project context (replace your\_project\_name with the actual listed name of your project)

```bash
icav2 projects enter "<your_project_name>"
```

3. List the pipelines in your project with

```bash
icav2 projectpipelines list
```

4. List the analyses inputs with the pipeline uuid, (not with the pipeline name).

```bash
icav2 projectpipelines input <your_pipeline_uuid>
```

### Minimal Example

JSON Pipelines are started with the following command:

```bash
icav2 projectpipelines start nextflowjson <your_pipeline_uuid> --pipeline_parameters
```

To retrieve the Platform Core file IDs for your input files, use:

```bash
icav2 projectdata list --file-name "<my_file_name>"
```

If you do not know the exact filename, you can search for files in your project with the command

```bash
icav2 projectdata list --file-name <part_of_the_filename> --match-mode fuzzy.
```

The command below launches a germline analysis with FASTQ inputs and the default reference, relying on form defaults for all other settings:

```bash
icav2 projectpipelines start nextflowjson \
  <pipeline-id> \
  --user-reference "my-germline-run" \
  --storage-size medium \
  --field-data fastqs:<fastq-file-id-1>,<fastq-file-id-2> \
  --field reference:"hg38_alt_masked_graph_v6"
```

### Key CLI Parameters

<table><thead><tr><th width="132.06640625">Field ID</th><th width="265.7109375">Example Value</th><th>Notes</th></tr></thead><tbody><tr><td>fastqs</td><td>&#x3C;file-id></td><td>Provide Platform Core file IDs for FASTQ inputs.</td></tr><tr><td>reference</td><td>"hg38_alt_masked_graph_v6"</td><td>Expand the drop-down in the UI for available values.</td></tr></tbody></table>

Omitted fields with defaults (e.g., enable\_map\_align, enable\_variant\_caller, enable\_cnv, enable\_sv, output\_format, enable\_dragen\_reports) are automatically applied from the form definition.

{% hint style="info" %}
The `icav2 projectpipelines input` command does not necessarily return all available fields. To discover the full set of available parameters, view the input form JSON from the pipeline's UI page in Platform Core.
{% endhint %}

{% hint style="info" %}
Some older or less commonly used pipelines use an XML-based input definition rather than `nextflowjson`. To launch these pipelines via the CLI, use the `nextflow` subcommand instead. Run `icav2 projectpipelines start nextflow --help` for usage details, as the parameter conventions differ. One key difference is that XML pipelines take a `ref_tar` input for the reference, where the user must provide the reference hash table as a file included in the DRAGEN bundle. See [Analysis Settings](#analysis-settings) for more details.
{% endhint %}

## Monitoring and Viewing Outputs

### Monitoring Analysis Status

1. Navigate to **Projects > your\_project > Flow > Analyses**.
2. Click the refresh button to update the status.
3. Click on a run to view details. The **Details** tab shows configuration, the **Nextflow execution** tab shows workflow progress, and the **Steps** tab shows logs (enable "Show technical steps" for additional log files).

The analysis status can also be monitored via the CLI:

```bash
icav2 projectanalyses get <id>
```

The `id` corresponds to the `id` field returned in the `projectpipelines start` command.

For more details on analysis states, see [Analysis Lifecycle](/project/p-flow/f-analyses#lifecycle).

{% hint style="info" %}
If the analysis failed, look at the [Debugging section](#debugging) to figure out what to do.
{% endhint %}

### Viewing Outputs

Analysis outputs can be viewed by navigating to the [analysis](/project/p-flow/f-analyses#starting-analyses) page in the GUI.

#### Report Tab

Most DRAGEN pipelines show an analysis report in the **Report** tab, unless it is disabled. The left-hand panel contains a **Summary** section with the overall `summary.html` (previously named `report.html`), as well as a **Samples** section listing individual per-sample reports. Selecting the summary report displays an interactive DRAGEN Reports page with tabs for key metrics, such as Summary, Enrichment, Trimmer, QC, Mapping, Coverage, and Variants. Selecting a sample report shows the same breakdowns for that sample.

<figure><img src="/files/nVaqrTRdodMucV4htE2N" alt=""><figcaption></figcaption></figure>

#### Output Files Tab

The **Output files** tab lists all files produced by the analysis. Smaller files can be downloaded directly from the browser, while larger files such as BAMs and VCFs should be downloaded via the CLI.

**Output JSON**

The output includes an `output.json` file with two top-level sections:

* **summary** — Counts of completed, failed, and total samples for the run.
* **samples** — A per-sample map keyed by sample name. Each entry includes the sample's processing status and analysis info, such as the reference genome and other DRAGEN options used.

This file is useful for reproducing the analysis, auditing the parameters that were applied, or programmatically checking which samples succeeded or failed.

**Analysis File Outputs**

Typical analysis file outputs might include:

* Alignment files (BAM/CRAM) with indexes
* VCF/GVCF files for small variants, CNVs, SVs, and STRs, where applicable and enabled
* Targeted caller reports (if enabled)
* QC metrics and coverage reports (if enabled)
* `summary.html` (if DRAGEN Reports is enabled; previously named `report.html`)

{% hint style="info" %}
If you have failed samples, you may notice that they do not appear in the report, have no output files, or have a status of "Failed" in the output.json. Refer to the [Debugging](#debugging) section for how to debug failed samples.
{% endhint %}

## Debugging

When you encounter a failed analysis, there are a few things to look for. The "Error" field on the main analysis UI page will, most of the time, give you a hint about the kind of error encountered. For multi-sample analyses, the output.json gives you summary statuses and the status for each sample. If the information above is insufficient, you can dig deeper into the process and pipeline runner logs.

### Finding the Failing Process

After identifying a failed analysis in **Projects > your\_project > Flow > Analyses**, navigate to the **Steps** tab of the analysis. A failing process will be marked with a non-zero exit code.

<figure><img src="/files/055gjlIa750zaACJWl0u" alt=""><figcaption></figcaption></figure>

### Finding the DRAGEN Command

Knowing the exact DRAGEN command that was executed is useful for debugging as well as for reproducing an analysis outside of the pipeline.

For DRAGEN processes, executed commands are logged in the **stderr**. DRAGEN commands start with `/opt/edico/bin/dragen`. The command can also be found in the **stdout** with the format:

```
Command Line: /opt/edico/bin/dragen ...
```

{% hint style="info" %}
Multiple samples may run in a single process depending on the `samples_per_node` input. A failed sample may not terminate the process, so failures can appear in the middle of the stdout log rather than at the end.
{% endhint %}

### Finding the Pipeline Runner Log

If no failing processes are visible, click the **Show technical steps** checkbox to reveal additional steps, including the pipeline runner stdout. Expand the `pipeline_runner.0` stdout to see Nextflow's own log messages, which will indicate which process failed and why.

<figure><img src="/files/89yJmDCrgecV9kUIHPyo" alt=""><figcaption></figcaption></figure>

## Analysis Settings

For most DRAGEN pipelines, the defaults are a good place to start. For more information on any of the input parameters beyond the parameter description, refer to the [DRAGEN User Guide](https://help.dragen.illumina.com/).

For a given published pipeline version, using the same set of parameters with a set of input data will give identical results, ensuring analyses are reproducible. Across different DRAGEN and/or pipeline versions, however, results may differ due to algorithmic improvements or the addition or removal of features. The best way to ensure the analysis is performed as similarly as possible across different versions is to check the DRAGEN options, which can be found in the [output.json](#output-files-tab) or the stdout of the DRAGEN process. Refer to the [Debugging](#debugging) section for more details on where to locate the stdout.

### Common Fields

**Input FASTQs / ORAs** — Pipelines will attempt to parse sample IDs from FASTQ filenames. To ensure that files are matched to the correct sample ID, users may optionally supply a FASTQ list to specify the structure. See the description in the input field for more details.

**samples\_per\_node** — This setting can be tweaked to optimize analysis runtime. For WGS samples, it is recommended to keep this at 1 sample per node. For exome or smaller panel samples, users can set it to 5 or higher.

**Storage Size** — Select a storage size equivalent to 2x the size of your input FASTQs (assuming BAM outputs).

**ref\_tar** — When supplying a custom reference to a JSON pipeline or any reference to an XML pipeline, ensure that the DRAGEN hash table version matches the DRAGEN version used by the pipeline. A mismatched hash table will cause the analysis to fail.

## Additional Resources

* [DRAGEN User Guide](https://help.dragen.illumina.com/)
* [DRAGEN Release Notes](https://support.illumina.com/sequencing/sequencing_software/dragen-bio-it-platform/downloads.html)
* [Platform Core End-to-End Tutorial](/tutorials/end-to-end-1)
* [Platform Core Pricing](/reference/r-pricing)




---

[Next Page](/llms-full.txt/1)

