JBrowse2: Difference between revisions
Created page with "{{DISPLAYTITLE:JBrowse 2}} JBrowse 2 is a genome browser that runs in your web browser. On Anunna you can start it from Open OnDemand, and it will show the data you already have on the cluster. This page explains what the JBrowse 2 app is, what each option on its form means, and which kinds of files it can show. It does not explain how to use JBrowse 2 itself. For that, see the [https://jbrowse.org/jb2/docs/ JBrowse 2 documentation]. == Introduction == === What JBrow..." |
mNo edit summary |
||
| (2 intermediate revisions by the same user not shown) | |||
| Line 1: | Line 1: | ||
{{DISPLAYTITLE:JBrowse 2}} | {{DISPLAYTITLE:JBrowse 2}} | ||
JBrowse 2 is a genome browser that runs in your web browser. On Anunna you can start it from | JBrowse 2 is a genome browser that runs in your web browser. On Anunna you can start it from the [https://apps.anunna.wur.nl Anunna Apps Portal], and it will show the data you already have on the cluster. | ||
This page explains what the JBrowse 2 app is, what each option on its form means, and which kinds of files it can show. It does not explain how to use JBrowse 2 itself. For that, see the [https://jbrowse.org/jb2/docs/ JBrowse 2 documentation]. | This page explains what the JBrowse 2 app is, what each option on its form means, and which kinds of files it can show. It does not explain how to use JBrowse 2 itself. For that, see the [https://jbrowse.org/jb2/docs/ JBrowse 2 documentation]. | ||
Latest revision as of 08:57, 1 October 2026
JBrowse 2 is a genome browser that runs in your web browser. On Anunna you can start it from the Anunna Apps Portal, and it will show the data you already have on the cluster.
This page explains what the JBrowse 2 app is, what each option on its form means, and which kinds of files it can show. It does not explain how to use JBrowse 2 itself. For that, see the JBrowse 2 documentation.
Introduction
What JBrowse 2 is
A genome browser works a bit like a map application, but for a genome. The reference genome is the map. Your data is drawn on top of it, as tracks: one track per file. You can zoom out to see a whole chromosome, or zoom in to see single bases.
JBrowse 2 can show, for example:
- sequencing reads aligned to the genome
- variants, such as SNPs
- coverage and other signals along the genome
- genes and other annotations
Why run it on Anunna
Sequencing data is large. A single alignment file can be many gigabytes, and you would not want to download it just to look at one gene.
With this app, you do not have to. JBrowse 2 runs on a compute node, right next to your data on Lustre. When you look at a region, only the small part of each file that covers that region is read. Nothing is copied.
Your files are also safe. The app can only read your data. It never changes, moves or adds anything in your folders.
Starting the app
Go to the Anunna Apps Portal, and find the JBrowse 2 tile under All Applications. Click on the button or selected it from the drop down, this will lead you to a form. Fill in the form, which is explained below, and click Launch.
Your session waits in the queue like any other job. Once it is running, a Connect to JBrowse 2 button appears. Click it, and JBrowse 2 opens in a new browser tab with your data already loaded.
The form
The form asks for two things about your data, the reference genome and a data folder, and a few things about the job itself.
| Option | What it means | Default |
|---|---|---|
| Maximum number of hours | How long the session may run. It stops after this time, even if you are still using it. Choose between 1 and 8 hours. | 4 |
| Reference genome (indexed FASTA) | The genome your data is aligned to. Required. | — |
| Data folder to browse | The folder that holds your data files. Optional. | the folder of the reference |
| Number of CPU cores | How many cores the session gets. | 2 |
| Total memory | How much memory the session gets. | 8 GB |
| Comment | A note for your job, for example a project number. | empty |
| SLURM options | Extra options for the job scheduler. Leave this empty if you are not sure. | empty |
| Email when started | Sends you an email when the session is ready. | off |
Reference genome
This is the most important choice. The reference is the "map" that everything else is drawn on, so all your data must be aligned to this same reference. If you choose a different version of the genome than the one your reads were aligned to, the tracks will not line up.
Use the browse button next to the field and choose a FASTA file. The file must:
- end in
.fa,.fastaor.fna, or in.fa.gz(and so on) if it is compressed withbgzip - have an index file next to it, with the same name plus
.fai, for examplegenome.faandgenome.fa.fai - if it is compressed, also have a
.gzifile next to it
If the index is missing, the session stops right away and tells you so. See Preparing your files for how to make one.
Data folder
This is the folder with the files you want to see. Use the browse button next to the field and choose a folder, not a file.
The app looks through this folder and its subfolders, up to four levels deep. Every file it can show becomes a track. Tracks are grouped by the subfolder they were found in, so a tidy folder gives you a tidy track list.
You can leave this field empty. The app then uses the folder that holds your reference genome. That is handy when the reference and the data are kept together.
A few tips:
- Choose the folder that holds the data you want to see, not a very large folder above it. The app adds at most 500 tracks. If a folder has more files than that, the rest are left out.
- Hidden files and folders, whose names start with a dot, are skipped.
- Subfolders that are only a link (a symlink) to somewhere else are not searched. Choose the real folder instead.
Cores and memory
A genome browser needs very little. JBrowse 2 does its drawing in your own web browser. The compute node only reads your files and sends the pieces you look at. The defaults, 2 cores and 8 GB, are enough for almost everyone. Asking for more will not make it faster, and it can make your session wait longer in the queue.
Supported data
The app can show these kinds of files. Most of them need an index file next to them. The index is a small file that tells JBrowse 2 where each part of the genome is stored in the big file, so it can jump straight to the region you are looking at.
| Data | File name ends in | Index needed next to it | Shown as |
|---|---|---|---|
| Aligned reads (BAM) | .bam |
.bai or .csi |
reads |
| Aligned reads (CRAM) | .cram |
.crai |
reads |
| Variants (VCF) | .vcf.gz |
.tbi or .csi |
variants |
| Annotation (GFF3) | .gff3.gz or .gff.gz |
.tbi or .csi |
features, such as genes |
| Regions (BED) | .bed.gz |
.tbi or .csi |
features |
| Signal (BigWig) | .bw or .bigwig |
none | a graph, such as coverage |
| Regions (BigBed) | .bb or .bigbed |
none | features |
Some things to know:
- Files without an index are skipped. The session still starts, but that file will not be in the track list.
- VCF, GFF3 and BED files must be compressed with
bgzip. A plain.vcf,.gff3or.bedfile is not shown. Note thatbgzipis not the same as normalgzip: a file compressed withgzipcannot be indexed. - CRAM files use the reference you chose to rebuild the reads. So for CRAM, choose the same reference genome the CRAM file was made with.
- Other file types in the folder, such as FASTQ or plain text, are simply ignored.
Preparing your files
If your files are not indexed yet, you need to do this once. After that, the index files stay next to your data and you never need to do it again.
The tools you need are in the same module the app uses. Please do this in a job, not on the login node. A small job script is enough:
#!/bin/bash
#SBATCH --job-name=index-for-jbrowse
#SBATCH --time=1:00:00
#SBATCH --mem=4G
#SBATCH --cpus-per-task=1
module load 2025 JBrowse2/4.3.0-GCC-14.2.0
# Go to the folder that holds your files
cd /lustre/path/to/your/data
# The reference genome: makes genome.fa.fai
samtools faidx genome.fa
# Aligned reads: makes sample1.bam.bai
samtools index sample1.bam
# Variants: compress, then index. Makes variants.vcf.gz and variants.vcf.gz.tbi
bgzip variants.vcf
tabix -p vcf variants.vcf.gz
Save this as, for example, index.sh, change the folder and file names to your own, and submit it with sbatch index.sh.
Use the line that matches each of your files:
| File | Command |
|---|---|
| Reference genome | samtools faidx genome.fa
|
| BAM or CRAM | samtools index sample1.bam
|
| VCF | bgzip variants.vcf, then tabix -p vcf variants.vcf.gz
|
| GFF3 | bgzip genes.gff3, then tabix -p gff genes.gff3.gz
|
| BED | bgzip regions.bed, then tabix -p bed regions.bed.gz
|
Indexing only works on files that are sorted by position along the genome. Alignment files made by most pipelines already are. If samtools index or tabix complains that the file is not sorted, sort it first, for example with samtools sort for a BAM file.
If a file is missing from the track list
When the session starts, the app writes a short report in the session's output.log. You can open it from the session card in Open OnDemand, through the link to the session folder. The report lists every track that was added, and every file that was skipped, with the reason.
The usual reasons are:
- the index file is missing, or has a different name than the data file
- the file ends in
.vcf,.gff3or.bedinstead of.vcf.gzand so on - the file is more than four folders below the data folder you chose
- the folder had more than 500 files the app could show
If the session does not start at all, the same log tells you why. Most often the reference genome has no .fai index, or the path you chose is not a FASTA file.