If you get an error "Unable to generate credentials from the objectstore as the requested path is too long." from AWS when requesting temporary credentials, then the path should be shortened.
You can truncate the sample name and user reference or use advanced output mapping in the API which avoids generating the long folders and creates output in the targetPath-defined location.
"analysisOutput": [
{
"sourcePath": "out",
"type": "FOLDER",
"targetProjectId": "enter_your_target_project_id",
"targetPath": "/enter_your_target_folder/"
}
]Before copy and move operations are executed on your own S3 storage, a test is performed to verify the necessary operational rights. If an issue is encountered, the copy/move action is aborted and the test files can be left behind due to missing rights. These files can safely be manually deleted from your S3 console.
Non-indexed folders () are designed for optimal performance in situations where no file actions are needed. They serve as fast storage in situations like temporary analysis file storage where you don't need access or searches via the GUI to individual files or subfolders within the folder. Think of a non-indexed folder as a data container. You can access the container which contains all the data, but you can not access the individual data files within the container from the GUI. As non-indexed folders contain data, they count towards your total project storage.
You can see the size of a non-indexed folder as part of the data details screen (Projects > your_project > Data > your_non-indexed_data > Data details tab) and in the data view (Projects > your_project > Data).
There can be a noticeable delay before the size of a non-indexed folder is updated after changes because of how the data is handled.
The GUI considers non-indexed folders as a single object. You can access the contents from a non-indexed folder
as Analysis input/output
in Bench
via the API
You can verify the integrity of the data by comparing the hash which is usually (with some exceptions) an MD5 (Message Digest Algorithm 5) checksum. This is a common cryptographic hash function that generates a fixed-size, 128-bit hash value from any input data. This hash value is unique to the content of the data, meaning even a slight change in the data will result in a significantly different MD5 checksum. AWS S3 calculates this checksum when data is uploaded and stores it in the ETag (Entity tag).
For files smaller than 16 MB, you can directly retrieve the MD5 checksum using our API endpoints. Make an API GET call to the https://ica.illumina.com/ica/rest/api/projects/{projectId}/data/{dataId} endpoint specifying the data Id you want to check and the corresponding project ID. The response you receive will be in JSON format, containing various file metadata. Within the JSON response, look for the objectETag field. This value is the MD5 checksum for the file you have queried. You can compare this checksum with the one you compute locally to ensure file integrity.
This ETag does not change and can be used as a file integrity check even when that file is archived, unarchived and/or copied to another location. Changes to the metadata have no impact on the ETag
For larger files, the process is different due to computation limitations. In these cases, we recommend using a dedicated pipeline on our platform to explicitly calculate the MD5 checksum. Below you can find both a main.nf file and the corresponding XML for a possible Nextflow pipeline to calculate the MD5 checksum for FASTQ files.
Since there is a storage cost associated with the data in your projects, it is good practice to regularly check how much cost is being generated by your projects and evaluate which data can be removed from cloud storage. The instructions provided here will help you quickly determine which data is generating the highest storage costs.
To see how much storage costs are currently being generated for your tenant, you can look at the usage explorer at https://platform.illumina.com/usage/ or from within Platform Core, navigate to the 9-dot symbol (
) in the top right next to your name and choose the usage explorer from the menu.
From the usage explorer overview screen, you can see below the graphical representation which projects are incurring the highest storage costs.

When you have determined which projects are incurring the largest storage costs, you can find out which files within that project are taking up the most space. To find the largest files in your project,
Go to Projects > your_project > Data and switch to list view with the () icon left above your files.
Use the column icon () top right to add the size column to your view. You can drag the size column to the desired position in your list view or use the move left and use right options which appear when you select the three vertical dots.
Select Sort descending to show the largest files first.
Once you have the list sorted like this, you can evaluate if those large files are still needed, if the can be sent to archive (manage > archive) or if they can be deleted (manage > delete).
Uploading Data
API Bench Analysis
Use non-indexed folders as normal folders for Analysis runs and bench. Different methods are available with the API such as creating temporary credentials to upload data to S3 or using /api/projects/{projectId}/data:createFileWithUploadUrl
Downloading Data
Yes
Use non-indexed folders as normal folders for Analysis runs and bench. Use temporary credentials to list and download data with the API.
Analysis Input/Output
Yes
Non-indexed files can be used as input for an analysis and the non-indexed folder can be used as output location. You will not be able to view the contents of the input and output in the analysis details screen.
Bench
Yes
Non-indexed folders can be used in Bench and the output from Bench can be written to non-indexed folders. Non-indexed folders are accessible across Bench workspaces within a project.
Viewing
No
The folder is a single object, you can not view the contents.
Linking
Yes
You cannot see non-indexed folder contents.
Copying
No
Prohibited to prevent storage issues.
Moving
No
Prohibited to prevent storage issues.
Managing tags
No
You cannot see non-indexed folder contents.
Managing format
No
You cannot see non-indexed folder contents.
Use as Reference Data
No
You cannot see non-indexed folder contents.
Creation
Yes
You can create non-indexed folders at Projects > your_project > Data > Manage > Create non-indexed folder. or with the /api/projects/{projectId}/data:createNonIndexedFolder endpoint
Deletion
Yes
You can delete non-indexed folders by selecting them at Projects > your_project > Data > select the folder > Manage > Delete.
or with the /api/projects/{projectId}/data/{dataId}:delete endpoint


nextflow.enable.dsl = 2
process md5sum {
container "public.ecr.aws/lts/ubuntu:22.04"
pod annotation: 'scheduler.illumina.com/presetSize', value: 'standard-small'
input:
file txt
output:
stdout emit: result
path '*', emit: output
publishDir "out", mode: 'symlink'
script:
txt_file_name = txt.getName()
id = txt_file_name.takeWhile { it != '.'}
"""
set -ex
echo "File: $txt_file_name"
echo "Sample: $id"
md5sum ${txt} > ${id}_md5.txt
"""
}
workflow {
txt_ch = Channel.fromPath(params.in)
txt_ch.view()
md5sum(txt_ch).result.view()
}<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<pd:pipeline xmlns:pd="xsd://www.illumina.com/ica/cp/pipelinedefinition">
<pd:dataInputs>
<pd:dataInput code="in" format="FASTQ" type="FILE" required="true" multiValue="true">
<pd:label>Input</pd:label>
<pd:description>FASTQ files input</pd:description>
</pd:dataInput>
</pd:dataInputs>
<pd:steps/>
</pd:pipeline>The Data inventory provides access to the files and folders stored in the project or linked to the project. Here, you can perform searches and data management operations such as moving, copying, deleting and (un)archiving.
See also which are a special form of data storage optimised for fast processing.
Platform Core supports UTF-8 characters in file and folder names for data. Please follow the guidelines detailed below. (For more information about recommended approaches to file naming that can be applicable across platforms, please refer to the AWS S3 documentation.)
Folders and files cannot be renamed after they have been created. To rename a folder, you will need to create a new folder with the desired name, move the contents from the original folder into the new one, and then delete the original folder. Please see the section for more information.
See the list of supported Data Formats
When adding data to Platform Core, prioritize data privacy. Whether you're using storage configurations like AWS S3 or performing Platform Core data uploads, careful management of data access needs to be considered. When setting up cloud storage, confirm that configuration settings prevent unauthorized access. Always verify that uploads are free of unintended data to avoid privacy breaches. For more detailed information, refer to the Platform Core Security and Compliance section.
See Data Integrity
On the Projects > your_project > Data page, you can view file information and preview files.
You can switch between folder view and flat view with the icons at the left.
Folder view shows the navigation structure and only the files and folders in the current folder. Searches in folder view will be performed on the current folder and all subfolders of the current folder.
Flat view shows a list of all files and folders within the current project. When you perform searches in flat view, all data of your project will be considered.
Tree view shows the navigation structure and only the files and folders in the current folder. Searches in tree view will be performed on the current folder and all subfolders of the current folder.
List view shows all files and folders within the current project. When you perform searches in list view, all data of your project will be considered.
To view file details click on the filename to see the file details.
Run input tags identifies the last 100 pipelines which used this file as input.
Run output tags identifies the pipeline which created the file.
Connector tags show if the file was added via browser upload or connector.
To view file contents, select the checkbox at the beginning of the line and then select View from the top menu. Alternatively, you can first click on the filename to see the details and then click the view tab to preview the file.
If your data is the result of an analysis, you can find the analysis which created it at Projects > your_project > Data > your_data > view > Data details tab > Source analysis. Clicking the link here will open the analysis.
To see the ongoing actions (copying and moving) on data in the data overview (Projects > your_project > Data), add the ongoing actions column from the column list if it is not present yet. You can also consult the data detail view for ongoing actions by clicking on the data in the overview. When clicking on an ongoing action itself, the data job details of the most recent created data job are shown.
If you open a folder by clicking it, you can see the folder details link at the top right. This will open the details screen where you can consult the folder size and number of files in that folder, the owning project, ongoing actions and folder id. You can also download the folder and all contents here with the download button.
To help navigate between folders in flat view, you can use the "path' column in the data view which will open the folder containing the selected file in tree view. If you want to go further up the folder path, you can use the folder structure above the file view. If the path is not visible, you can add it with the three-columns symbol next to the filter symbol.
To quickly find data, use the search dialog at the top right. Search is performed with automatic wildcards before and after the search text. You can use * as additional wildcard, for example b*n will match bunny as it is interpreted as *b*n*.
You can search on the file name, the path (/folder/subfolder) and tags.
When Secondary Data is added to a data record, those secondary data records are mounted in the same parent folder path as the primary data file when the primary data file is provided as an input to a pipeline. Secondary data is intended to work with the CWL feature. This is commonly used with genomic data such as BAM files with companion BAM index files.
You can create hyperlinks to data to quickly share it with the following syntax:
You can export the list of data which you see in the overview as a CSV, JSON, or excel file.
Select one or more files to export at Projects > your_project > Data.
Select Export at the bottom of the screen.
Choose between the following export options:
Single files can be downloaded directly from within the UI.
Select the checkbox next to the file which you want to download, followed by Download > Browser Download > Download.
You can also download files from their details screen. Click on the file name and select Download at the bottom of the screen. Depending on the size of your file, it may take some time to load the file contents.
You can trigger an asynchronous download via service connector using the Schedule for Download button with one or more files selected.
Select a file or files to download.
Select Download > Schedule download (for files or folders). This will display a list of all available connectors.
Select a connector and optionally, enter your email address if you want to be notified of download completion, and then select Download.
You can view the progress of the download or abort the scheduled download on the page for the project.
Uploading data to the platform makes it available to analysis workflows and tools.
To upload data manually via the drag-and-drop interface in the platform UI, go to Projects > your_project > Data and either
Drag a file from your system into the Choose a file or drag it here box.
Select the Choose a file or drag it here box, and then choose a file. Select Open to upload the file.
Your files are added to the Data page with status partial during upload and become available when upload completes.
For instructions on uploading/downloading data via CLI, see .
You can copy data from your project to a different folder within the same project or you can copy data from another project to your current project, provided you have the necessary access rights.
You can copy data from a subfolder to a higher-level folder to move data up one or more levels (folder/destination/source). You can not copy data from the source folder onto itself or onto a subfolder of the source folder as this would result in a loop.
The person copying the data must have the following rights:
The following restrictions apply when copying data:
Go to the destination project for your data copy and proceed to Projects > your_project > Data > Manage > Copy From.
Optionally, use the filters or search with the search box for the desired data.
Select the data (individual files or folders with data) you want to copy.
INITIALIZED
WAITING_FOR_RESOURCES
RUNNING
STOPPED - When choosing to stop the batch job.
To see the ongoing actions on data in the data overview (Projects > your_project > Data), you can add the ongoing actions column from the column list with the three column symbol at the top right, next to the filter funnel. You can also consult the data detail view for ongoing actions by clicking on the data in the overview.
You can move data within a project or between different projects to which you have access. If your browser allows notifications, a pop-up will appear when the move is completed.
Move From is used when you are in the destination location.
Move To is used when you are in the source location.
Before moving the data, pre-checks are performed to verify that the data can be moved and no currently running operations are being performed on the folder. Conflicting jobs and missing permissions will be reported.
Once the move has started, no other operation must be performed on the data being moved to avoid potential data loss or duplication. When modifying data at the source or destination during a move process, incomplete data transfers may occur with duplicate folders and files with different identifiers.. You can manually transfer any remaining data and delete duplicate files and folders afterward.
There are a number of rights and restrictions related to data move as this will delete the data in the source location.
1000 Maximum Items: Up to 1000 items per move. Items include files and folders. Folders with subfolders and subfiles still count as one item.
Naming Conflicts: Cannot move to a destination with existing files/folders of the same name.
Linked Data Restrictions: Cannot move linked data move data to linked data.
Move Data From is used when you are in the destination location.
Navigate to Projects > your_project > Data > your_destination_location > Manage > Move From.
Select the files and folders which you want to move.
Select the Move button.
Move Data To is used when you are in the source location. You will need to select the data you want to move from to current location and the destination to move it to.
Navigate to Projects > your_project > Data > your_source_location.
Select the files and folders which you want to move.
Select to Projects > your_project > Data > your_source_location > Manage > Move To.
INITIALIZED
WAITING_FOR_RESOURCES
RUNNING
STOPPED - When choosing to stop the batch job.
To see the ongoing actions on data in the data overview (Projects > your_project > Data), add the ongoing actions column from the column list with the three column symbol at the top right, next to the filter funnel. You can also consult the data detail view for ongoing actions by clicking on the data in the overview.
To manually archive or delete files:
Select the checkbox next to the file or files to delete or archive.
Select Manage, and then select one of the following options:
Archive — Move the file or files to long-term storage (event code ICA_DATA_110).
When attempting concurrent archiving or unarchiving of the same file, a message will inform you to wait for the currently running (un)archiving to finish first.
To archive or delete files programmatically, you can use Platform Core's API endpoints:
the file's information.
Modify the dates of the file to be deleted/archived.
the updated information back in Platform Core.
Data linking creates a dynamic read-only view to the source data. You can use data linking to get access to data without running the risk of modifying the source material and to share data between projects. Linking ensures changes to the source data are immediately visible and no additional storage is required. You can recognise linked data by the green color and see the owning project as part of the details.
Since this is read-only access, you cannot perform actions such as deleting, adding, moving or (un)archiving on linked data as these actions require write access.
Select Projects > your_project > Data > Manage, and then select Link.
To view data by project, select the funnel symbol, and then select Owning Project. If you know to which project the data is linked to, you can choose to filter on linked projects. If you click on a folder, the folder will open so you can access the files, if you click on a file, the file details will be opened.
Select the checkbox next to the file or files to add.
Your files will be added and visible in the Data page.
If you link a folder instead of individual files, a warning is displayed indicating that, depending on the size of the folder, linking may take considerable time. The linking process will run in the background and the progress can be monitored on the Projects > your_project > activity > Batch Jobs screen. From here you can see more details such as how many files have already been linked, by clicking the batch job.
To unlink the data, go to the root level of your project and select the linked folder or, if you have linked individual files separately, then you can select those linked files (limited to 100 at a time) and select Manage > Unlink. The progress can be monitored at Projects > your_project > Activity > Batch Jobs.
AnalysisID
At YourProject > Flow > Analyses > YourAnalysis > ID
To export only the columns which are currently shown in your view, select Visible columns as the Columns to export option, otherwise, choose All columns.
Select the export format. (CSV/JSON/Excel)
Select which action to take if the data already exists (overwrite existing data, don't copy or keep both the original and the new copy by appending a version number to the copied data).
Select Copy to copy the data to your project. You can see the progress in Projects > your_project > Activity > Batch Jobs and if your browser permits it, a pop-up message will be displayed when the copy process completes.
SUCCEEDED - All files and folders are copied.
PARTIALLY_SUCCEEDED - Some files and folders could be copied, but not all. Partially succeeded will typically occur when files were being modified or unavailable while the copy process was running.
FAILED - None of the files and folders could be copied.
Self Move: Folders cannot be moved to themselves.
In-Transit Data: Cannot move data that is being moved.
Region Restrictions: No cross-region moves allowed.
Project Constraints: No moves from externally-managed projects or externally-managed data.
Status Requirement: Data must be in status available.
Ownership: Data must be owned by the user's tenant for cross-project moves.
Destination Default: If no target folder is selected, data moves to the root folder of the target project.
Select the Move button.
SUCCEEDED - All files and folders are moved.
PARTIALLY_SUCCEEDED - Some files and folders could be moved, but not all. Partially succeeded will typically occur when files were being modified or unavailable while the move process was running.
FAILED - None of the files and folders could be moved.
Delete — Remove the file completely (event code ICA_DATA_106).
Select Link.
a-z
A-Z
Special characters
Exclamation point !
Hyphen -
Underscore _
Period .
Asterisk *
Single quote '
Open parenthesis (
Closed parenthesis )
https://<ServerURL>/ica/link/project/<ProjectID>/data/<FolderID>https://<ServerURL>/ica/link/project/<ProjectID>/analysis/<AnalysisID>ServerURL
See browser address bar.
projectID
At YourProject > Details > URN > urn:ilmn:ica:project:ProjectID#MyProject
FolderID
At YourProject > Data > folder > folder details > ID
Within a project
Contributor rights
Upload and Download rights
Contributor rights
Upload and Download rights
Between different projects
Download rights
Viewer rights
Within a project
No linked data
No partial data
No archived data
No Linked data
Between different projects
Data sharing enabled
No partial data
No archived data
Within the same region
Within a project
Contributor rights
Contributor rights
Between different projects
Download rights
Contributor rights
Within a project
No linked data
No partial data
No archived data
No Linked data
Between different projects
Data sharing enabled
Data owned by user's tenant
No linked data
No partial data
No archived data
No externally managed projects
Within the same region
You cannot switch between folder and flat view when you are viewing search results. Clear the search with the clear search button, or the x in the search dialog first.
When you share the data view by sharing the link from your browser, filters and sorting is retained in links, so the recipient will see the same data and order.
Normal permission checks still apply with these links. If you try to follow a link to data to which you do not have access, you will be returned to the main project screen or login screen, depending on your permissions.
To prevent cost issues, you can not perform actions such as copying and moving data which would write data to the workspace when the project billing mode is set to tenant and the owning tenant of the folder is not the current user's tenant.
If you do not have a connector, create one and install it. You must then return to the file selection in step 1 to use it.
Do not close the Platform Core tab in your browser while data uploads.
Uploads via the UI are limited to 5TB and no more than 100 concurrent files at a time, but for practical and performance reasons, it is recommended to use the CLI or Service connector when uploading large amounts of data.
Copying large amounts of data can take considerable time. You can monitor the progress at Projects > your_project > Activity > Batch Jobs.
Data in the "Partial" or "Archived" state will be skipped during a copy job.
There is a difference in copy type behavior between copying files and folders. The behavior is designed for files and it is best practice to not copy folders if there already is a folder with the same name in the destination location.
Notes on copying data
Copying data comes with an additional storage cost as it will create a copy of the data.
Copying data from your own S3 storage requires additional configuration. See Connect AWS S3 Bucket and SSE-KMS Encryption.
On the command-line interface, the command to copy data is icav2 projectdata copy.
Before copy and move operations are executed on your own S3 storage, a test is performed to verify the necessary operational rights. This can result in temporary test files remaining (for example when is not correctly set up for a versioned bucket). These files can safely be manually deleted from your S3 console.
Changes to the date during move may cause the destination data to be unsynchronized between the object store (S3) and Platform Core. To address this, create a folder session on the destination directory's parent folder by using the following API steps: Create Folder Session and Complete Folder Session. Ensure that the move job is aborted before making the create and complete requests for the folder session.
Move jobs will fail if any data being moved is in the Partial or Archived state.
Moving large amounts of data can take considerable time. You can monitor the progress at Projects > your_project > Activity > Batch Jobs.
Moving large amounts of data can take considerable time. You can monitor the progress at Projects > your_project > Activity > Batch Jobs.
If you are only able to select your source project as the target data project, this may indicate that data sharing (Projects > your_project > Project Settings > Details > Data Sharing) is not enabled for your project or that you do not have have upload rights in other projects.
Before copy and move operations are executed on your own S3 storage, a test is performed to verify the necessary operational rights. This can result in temporary test files remaining (for example when IAM policy is not correctly set up for a versioned bucket). These files can safely be manually deleted from your S3 console.
The Python snippet below exemplifies the approach: it sets (or updates if set already) the time to be archived for a specific file:
import requests
import json
from config import PROJECT_ID, DATA_ID, API_KEY
url_get="https://ica.illumina.com/ica/rest/api/projects/" + PROJECT_ID + "/data/" + DATA_ID
# set the API get headers
headers = {
'X-API-Key': API_KEY,
'accept': 'application/vnd.illumina.v3+json'
}
# set the API put headers
headers_put = {
'X-API-Key': API_KEY,
'accept': 'application/vnd.illumina.v3+json',
'Content-Type': 'application/vnd.illumina.v3+json'
}
# Helper function to insert willBeArchivedAt after field named 'region'
def insert_after_region(details_dict, timestamp):
new_dict = {}
for k, v in details_dict.items():
new_dict[k] = v
if k == 'region':
new_dict['willBeArchivedAt'] = timestamp
if 'willBeArchivedAt' in details_dict:
new_dict['willBeArchivedAt'] = timestamp
return new_dict
# 1. Make the GET request
response = requests.get(url_get, headers=headers)
response_data = response.json()
# 2. Modify the JSON data
timestamp = "2024-01-26T12:00:04Z" # Replace with the provided timestamp
response_data['data']['details'] = insert_after_region(response_data['data']['details'], timestamp)
# 3. Make the PUT request
put_response = requests.put(url_get, data=json.dumps(response_data), headers=headers_put)
print(put_response.status_code)To delete a file at specific timepoint, the key 'willBeDeletedAt' should be added or changed using the API call. If running in the terminal, a successful run will finish with the message ‘200’. In the Platform Core UI, you can check the details of the file to see the updated values for ‘Time To Be Archived’ (willBeArchivedAt) or ‘Time To Be Deleted’ (willBeDeletedAt), as shown in the screenshot.
Linking data is only possible from the root folder of your destination project. The action is disabled in project subfolders.
Linking a parent folder after linking a file or subfolder will unlink the file or subfolder and link the parent folder. So root\linked_subfolder will become root\linked_parentfolder\linked_subfolder.
Initial linking can take considerable time when there is a large amount of source data. However, once the initial link is made, updates to the source data will be instantaneous. You can monitor the progress at Projects > your_project > activity > Batch Jobs.
Before Platform Core version v.2.29, when data was linked, a snapshot was created of the file and folder structure. These links created a read-only view of the data as it was at the time of linking, but did not propagate changes to the file and folder structure. If you want to use the advantages of the new way of linking with dynamic updates, unlink the data and relink it. Since snapshot linking has been deprecated, all new data linking done in Platform Core v.2.29 or later has dynamic content updates.
Filtering
To add filters, select the funnel/filter symbol at the top right, next to the search field.
Filters are reset when you exit the current screen.
Sorting
To sort data, select the three vertical dots in the column header on which you want to sort and chose ascending or descending.
Sorting is retained when you exit the current screen.
Displaying Columns
To change which columns are displayed, select the three columns symbol and select which columns should be shown.
In , you can see which files are externally controlled and which are ICA-managed by means of the “managed by” column.
In view, you can jump to the folder in which files are located with the "Path" column.
The displayed columns are retained when you exit the current screen.
Replace
Overwrites the existing data. Folders will copy their data in an existing folder with existing files. Existing files will be replaced when a file with the same name is copied and new files will be added. The remaining files in the target folder will remain unchanged.
Don't copy
The original files are kept. If you selected a folder, files that do not yet exist in the destination folder are added to it. Files that already exist at the destination are not copied over and the originals are kept.
Keep both
Files have a number appended to them if they already exist. If you copy folders, the folders are merged, with new files added to the destination folder and original files kept. New files with the same name get copied over into the folder with a number appended.






Upload rights
Contributor rights
No linked data
Within the same region
Upload rights
Viewer rights
No linked data
Within same region

