Jump to content

Main Page: Difference between revisions

From HPCwiki
Herre011 (talk | contribs)
mNo edit summary
 
(147 intermediate revisions by 16 users not shown)
Line 1: Line 1:
The Agrogenomics cluster is a [http://en.wikipedia.org/wiki/High-performance_computing High Performance Compute] (HPC) infrastructure hosted by [http://www.wageningenur.nl/nl/activiteit/Opening-High-Performance-Computing-cluster-HPC.htm Wageningen University & Research Centre]. It is open for use for all WUR research groups as well as other organizations, including companies, that have collaborative projects with WUR.  
Anunna is a [http://en.wikipedia.org/wiki/High-performance_computing High Performance Computing] (HPC) cluster hosted by [https://www.wur.nl/ Wageningen University & Research]. It is open to all WUR research groups, and to other organisations and companies running collaborative projects with WUR.


The Agrogenomics HPC was an initiative of the [http://www.breed4food.com/en/breed4food.htm Breed4Food] (B4F) consortium, consisting of the [[About_ABGC | Animal Breeding and Genomics Centre]] (WU-Animal Breeding and Genomics and Wageningen Livestock Research) and four major breeding companies: [http://www.cobb-vantress.com Cobb-Vantress], [https://www.crv4all.nl CRV], [http://www.hendrix-genetics.com Hendrix Genetics], and [http://www.topigs.com TOPIGS]. Currently, in addition to the original partners, the HPC (HPC-Ag) is used by other groups from Wageningen UR (Bioinformatics, Centre for Crop Systems Analysis, Environmental Sciences Group, and Plant Research International) and plant breeding industry (Rijk Zwaan).  
You reach the cluster over [[SSH Access|SSH]] at <code>login.anunna.wur.nl</code>, or through your browser at the [[Apps Portal]] (https://apps.anunna.wur.nl/).


== Rationale and Requirements for a new cluster ==
== [[About]] ==
[[File:Breed4food-logo.jpg|thumb|right|200px|The Breed4Food logo]]
* <h3>[[Mission and Governance]]</h3> What Anunna is for and who runs it
The Agrogenomics Cluster was originally conceived as being the 7th pillar of the [http://www.breed4food.com/en/show/Breed4Food-initiative-reinforces-the-Netherlands-position-as-an-innovative-country-in-animal-breeding-and-genomics.htm Breed4Food programme]. While the other six pillars revolve around specific research themes, the Cluster represents a joint infrastructure. The rationale behind the cluster is to enable the increasing computational needs in the field of genetics and genomics research, by creating a joint facility that will generate benefits of scale, thereby reducing cost. In addition, the joint infrastructure is intended to facilitate cross-organisational knowledge transfer. In that capacity, the HPC-Ag acts as a joint (virtual) laboratory where researchers - academic and applied - can benefit from each other's know-how. Lastly, the joint cluster, housed at Wageningen University campus, allows retaining vital and often confidential data sources in a controlled environment, something that cloud services such as Amazon Cloud or others usually can not guarantee.
* <h3>[[Roadmap]]</h3> How we decide what to work on
{{-}}
* <h3>[[Cluster Architecture Overview]]</h3> How the cluster is put together
* <h3>[[Compute_Hardware_Overview]]</h3> The node types and their hardware
* <h3>[[Storage Systems Overview]]</h3> The storage tiers at a glance
* <h3>[[Tariffs]]</h3> Costs of using Anunna
* <h3>[[Network & Security]]</h3> Interconnect and data-security posture
* <h3>[[History of the Cluster]]</h3> How Anunna came to be
* <h3>[[Sustainability]]</h3> Green-HPC policy and measures
* <h3>[[FAQ]]</h3> Frequently Asked Questions
* <h3>[[Policies and Terms of Use]]</h3> Who may use Anunna and on what terms
*


== Process of acquisition and financing ==
== [[Get Started]] ==
* <h3>[[Who Can Access?]]</h3> Eligibility
* <h3>[[Account Application Process]]</h3> How to request an account
* <h3>[[User Responsibilities]]</h3> What is expected of you as a user
* <h3>[[SSH Access]]</h3> Logging in over SSH to the cluster
* <h3>[[Apps Portal]]</h3> Using applications on Anunna through your web browser
* <h3>[[Workflow Migration from Laptop to HPC]]</h3> Moving your work to the cluster from your computer
* <h3>[[Grant_Support_%26_Acknowledgement|Grant support & Acknowledgement]]</h3> Guidance and standard text that can be used in grant proposals, research papers, theses and other publications.


[[File:Signing_CatAgro.png|thumb|left|300px|Petra Caessens, manager operations of CAT-AgroFood, signs the contract of the supplier on August 1st, 2013. Next to her Johan van Arendonk on behalf of Breed4Food.]]
== [[System Access]] ==
The Agrogenomics cluster was financed by [http://www.wageningenur.nl/en/Expertise-Services/Facilities/CATAgroFood-3/CATAgroFood-3/News-and-agenda/Show/CATAgroFood-invests-in-a-High-Performance-Computing-cluster.htm CATAgroFood]. The [[B4F_cluster#IT_Workgroup | IT-Workgroup]] formulated a set of requirements that in the end were best met by an offer from [http://www.dell.com/learn/nl/nl/rc1078544/hpcc Dell]. [http://www.clustervision.com ClusterVision] was responsible for installing the cluster at the Theia server centre of FB-ICT.
* <h3>[[Login Nodes]]</h3> The entry point you connect to
{{-}}
* <h3>[[Compute Nodes]]</h3> Where your jobs run, and how to reach them
* <h3>[[Filesystems]]</h3> The storage tiers
* <h3>[[Quotas]]</h3> Storage limits and how to check them
* <h3>[[Remote Access (VPN, Gateway)]]</h3> Reaching Anunna from off campus
* <h3>[[Data Transfer Methods]]</h3> Moving data to and from Anunna


== Architecture of the cluster ==
== [[Scheduling|Job Scheduling and Resource Management]] ==
[[Architecture_of_the_HPC | Main Article: Architecture of the Agrogenomics HPC]]
* <h3>[[Scheduler Overview (Slurm)]]</h3> How scheduling works
[[File:Cluster_scheme.png|thumb|right|600px|Schematic overview of the cluster.]]
* <h3>[[Partitions / Queues]]</h3> The available partitions
The new Agrogenomics HPC has a classic cluster architecture: state of the art Parallel File System (PSF), headnodes, compute nodes (of varying 'size'), all connected by superfast network connections (Infiniband). Implementation of the cluster will be done in stages. The initial stage includes a 600TB PFS, 48 slim nodes of 16 cores and 64GB RAM each, and 2 fat nodes of 64 cores and 1TB RAM each. The overall architecture, that include two head nodes in fall-over configuration and an infiniband network backbone, can be easily expanded by adding nodes and expanding the PFS. The cluster management software is designed to facilitate a heterogenous and evolving cluster.
* <h3>[[Choosing a node (constraints)]]</h3> Targeting particular hardware
{{-}}
* <h3>[[Batch Jobs]]</h3> Submitting a batch script
* <h3>[[Interactive Jobs]]</h3> An interactive shell on a compute node
* <h3>[[Array Jobs]]</h3> Running many similar jobs at once
* <h3>[[Monitoring Jobs]]</h3> Checking on your jobs
* <h3>[[Cancelling Jobs]]</h3> Stopping a job
* <h3>[[Reservations]]</h3> Reserving nodes
* <h3>[[Fair Use Policy]]</h3> Using shared resources considerately


== Housing at Theia ==
== [[Software]] ==
[[File:Map_Theia.png|thumb|left|200px|Location of Theia, just outside of Wageningen campus]]
* <h3>[[Software Overview]]</h3> The software landscape
The Agrogenomics Cluster is housed at one of two main server centres of WUR-FB-IT, near Wageningen Campus. The building (Theia)  may not look like much from the outside (used to function as potato storage) but inside is a modern server centre that includes, a.o., emergency power backup systems and automated fire extinguishers. Many of the server facilities provided by FB-ICT that are used on a daily basis by WUR personnel and students are located there, as is the Agrogenomics Cluster. Access to Theia is evidently highly restricted and can only be granted in the presence of a representative of FB-IT.
* <h3>[[Environment Modules]]</h3> Loading software with modules and buckets
{{-}}
* <h3>[[Installing Personal Software]]</h3> Installing into your own space
{| width="90%"
* <h3>[[Licensed Software]]</h3> Software that needs a licence
|- valign="top"
| width="10%" |


| width="30%" |
* <h3>Scripting languages</h3>
[[File:Cluster2_pic.png|thumb|left|220px|Some components of the cluster after unpacking.]]
** <h4>[[Python]]</h4> Python modules and environment managers
| width="70%" |
** <h4>[[R]]</h4> R modules, packages, and parallel R jobs
[[File:Cluster_pic.png|thumb|right|400px|The final configuration after installation.]]
** <h4>[[Julia]]</h4> Loading Julia, packages, and running it in a job
|}
{{-}}


== Management ==
* <h3>Containers</h3>
[[HPC_management | Main Article: HPC management]]
** <h4>[[Apptainer]]</h4> Portable, reproducible containers for HPC


Project Leader of the HPC is Stephen Janssen (Wageningen UR,FB-IT, Service Management). [[User:pollm001 | Koen Pollmann (Wageningen UR,FB-IT, Infrastructure)]] and [[User:dawes001 | Gwen Dawes (Wageningen UR, FB-IT, Infrastructure)]] are responsible for [[Maintenance_and_Management | Maintenance and Management]].
* <h3>MPI implementations</h3>
** <h4>[[OpenMPI]]</h4> The MPI library Anunna is built around
** <h4>[[IntelMPI]]</h4> Intel MPI, with the Intel compilers and MKL


== Access Policy ==
== [[Storage]] ==
[[Access_Policy | Main Article: Access Policy]]
* <h3>[[Storage Systems Overview]]</h3> The storage tiers at a glance
* <h3>[[Home Directory]]</h3> Your personal space
* <h3>[[Compute Storage]]</h3> The fast Lustre filesystem for active work
* <h3>[[Shared Storage]]</h3> Sharing data within a group
* <h3>[[Backup Policy]]</h3> What is backed up, and what is not
* <h3>[[Archival Storage]]</h3> Long-term storage on tape
* <h3>[[Quotas]]</h3> Storage limits
* <h3>[[Data Lifecycle Policy]]</h3> How data moves from active to archived to removed
* <h3>[[Data storage best practices|Data Storage Best Practices]]</h3> Keeping data safe, tidy, and cheap
* <h3>[[Data Transfer Best Practices]]</h3> Moving data efficiently and reliably
* <h3>[[Handling Sensitive Data (GDPR, etc.)]]</h3> Confidential and personal data


Access policy is still a work in progress. In principle, all staff and students of the five main partners will have access to the cluster. Access needs to be granted actively (by creation of an account on the cluster by FB-IT reagdring non WUR accounts). Use of resources is limited by the scheduler. Depending on availability of queues ('partitions') granted to a user, priority to the system's resources is regulated.
== Graphical Interface Applications ==
* <h3>[[Portal Overview]]</h3> What the portal is and how the dashboard is laid out
* <h3>[[Running GUI Applications]]</h3> The general way to launch and connect to a graphical application
* <h3>[[How to Launch a Desktop Session]]</h3> Start a full Linux desktop in your browser
* <h3>[[Jupyter]]</h3> Jupyter notebooks
* <h3>[[RStudio]]</h3> The RStudio IDE for R and Python
* <h3>[[File Browser]]</h3> Manage your files in the browser
* <h3>[[Shell Access]]</h3> A terminal in the browser
* <h3>[[Jobs Queue Overview]]</h3> See your jobs in the browser


== Users ==
== [[Training]] ==
* <h3>[[Training Materials]]</h3> Course slides and self-study resources
* <h3>[[Tutorials]]</h3> Hands-on tutorials
* <h3>[[Workshops]]</h3> Instructor-led courses and their dates


* [[List_of_users | List of users (alphabetical order)]]
== [[Support]] ==
* [[Mailinglist | Electronic mail discussion lists]]
* <h3>[[How to Get Help]]</h3> Who to contact and how
* <h3>[[Support Ticket System]]</h3> The WUR support portal
* <h3>[[Reporting Incidents]]</h3> Writing a good problem report
* <h3>[[Known Issues]]</h3> Current problems and workarounds
* <h3>[[Our User community]]</h3> The Anunna user community
* <h3>[[Maintenance Schedule]]</h3> Planned downtimes


== Using the HPC-Ag ==
== [[For PIs]] ==
=== Gaining access to the HPC-Ag ===
* <h3>[[Dos and Don'ts]]</h3> Good and bad practice at a glance
Access to the cluster and file transfer are done by [http://en.wikipedia.org/wiki/Secure_Shell ssh-based protocols].
* <h3>[[Managing Group Members]]</h3> Working with groups
* [[log_in_to_B4F_cluster | Logging into cluster using ssh and file transfer]]
* <h3>[[Storage Requests]]</h3> Requesting more storage
* <h3>[[Resource Allocation Requests]]</h3> Requesting larger or dedicated allocations
* <h3>[[External Collaborator Access]]</h3> Giving access to collaborators outside WUR
* <h3>[[Reporting Usage]]</h3> Tracking your group's usage and costs
* <h3>[[Grant Support & Acknowledgement]]</h3> Facility descriptions and support for funding applications


=== Cluster Management Software and Scheduler ===
== [[Workflows]] ==
The HPC-Ag uses Bright Cluster Manager software for overall cluster management, and Slurm as job scheduler.
* <h3>[[Workflow Migration from Laptop to HPC]]</h3> Moving your work to the cluster
* [[BCM_on_B4F_cluster | Monitor cluster status with BCM]]
* <h3>[[Reproducibility Guidelines]]</h3> Keeping your work reproducible
* [[SLURM_on_B4F_cluster | Submit jobs with Slurm]]
* <h3>[[Workflow Engines (Snakemake, Nextflow)]]</h3> Managing multi-step pipelines
* [[SLURM_Compare | Rosetta Stone of Workload Managers]]
* <h3>[[Debugging Jobs]]</h3> Working out why a job failed
* <h3>[[Checkpointing]]</h3> Saving and restarting long jobs
* <h3>[[Scheduled tasks (cron)|Scheduled Tasks (cron)]]</h3> Running recurring tasks with scrontab


=== Installation of software by users ===
* <h3>Parallel Workflows</h3>
** <h4>[[Workflows/Parallel-Computing|Parallel Computing]]</h4> What parallel computing means on Anunna, and which type fits your work
** <h4>[[Workflows/Serial|Serial]]</h4> One program on a single core — the baseline
** <h4>[[Workflows/Embarassinly-Parallel|Embarrassingly Parallel]]</h4> Running the same program many times over
** <h4>[[Workflows/Multi-threaded|Multi-threaded]]</h4> Many cores on one machine, within a single process
** <h4>[[Workflows/Multi-Process|Multi-Process]]</h4> One calculation spread across several machines


* [[Domain_specific_software_on_B4Fcluster_installation_by_users | Installing domain specific software: installation by users]]
== Quick links ==
* [[Setting local variables]]
* [[Installing_R_packages_locally | Installing R packages locally]]
* [[Setting_up_Python_virtualenv | Setting up and using a virtual environment for Python3 ]]


=== Installed software ===
* [[How to Get Help]] — contact the HPC team
 
* [[Maintenance Schedule]] — planned downtimes
* [[Globally_installed_software | Globally installed software]]
* [[Glossary of Terms]] — common HPC and Anunna terms explained
* [[ABGC_modules | ABGC specific modules]]
 
=== Being in control of Environment parameters ===
 
* [[Using_environment_modules | Using environment modules]]
* [[Setting local variables]]
* [[Setting_TMPDIR | Set a custom temporary directory location]]
* [[Installing_R_packages_locally | Installing R packages locally]]
* [[Setting_up_Python_virtualenv | Setting up and using a virtual environment for Python3 ]]
 
=== Controlling costs ===
 
* [[SACCT | using SACCT to see your costs]]
* [[get_my_bill | using the "get_my_bill" script to estimate costs]]
 
== Miscellaneous ==
* [[Bioinformatics_tips_tricks_workflows | Bioinformatics tips, tricks, and workflows]]
* [[Convert_between_MediaWiki_and_other_formats | Convert between MediaWiki format and other formats]]
* [[Manual GitLab | Create projects and add scrits]]
 
== See also ==
* [[Maintenance_and_Management | Maintenance and Management]]
* [[Mailinglist | Electronic mail discussion lists]]
* [[About_ABGC | About ABGC]]
* [[Computer_cluster | High Performance Computing @ABGC]]
* [[Lustre_PFS_layout | Lustre Parallel File System layout]]


== External links ==
== External links ==
{| width="90%"
|- valign="top"
| width="30%" |
* [http://www.breed4food.com/en/show/Breed4Food-initiative-reinforces-the-Netherlands-position-as-an-innovative-country-in-animal-breeding-and-genomics.htm Breed4Food programme]
* [http://www.wageningenur.nl/en/Expertise-Services/Facilities/CATAgroFood-3/CATAgroFood-3/Our-facilities/Show/High-Performance-Computing-Cluster-HPC.htm CATAgroFood offers a HPC facilty]
* [http://www.cobb-vantress.com Cobb-Vantress homepage]


| width="30%" |
* [https://www.wur.nl/en/Value-Creation-Cooperation/Facilities/Wageningen-Shared-Research-Facilities/Our-facilities/Show/High-Performance-Computing-Cluster-HPC-Anunna.htm Wageningen Shared Research Facilities — HPC]
* [https://www.crv4all.nl CRV homepage]
* [http://www.hendrix-genetics.com Hendrix Genetics homepage]
* [http://www.topigs.com TOPIGS homepage]
| width="30%" |
* [http://en.wikipedia.org/wiki/Scientific_Linux Scientific Linux]
* [http://en.wikipedia.org/wiki/Help:Cheatsheet Help with editing Wiki pages]
|}

Latest revision as of 06:49, 31 August 2026

Anunna is a High Performance Computing (HPC) cluster hosted by Wageningen University & Research. It is open to all WUR research groups, and to other organisations and companies running collaborative projects with WUR.

You reach the cluster over SSH at login.anunna.wur.nl, or through your browser at the Apps Portal (https://apps.anunna.wur.nl/).

  • Scripting languages

    • Python modules and environment managers
    • R modules, packages, and parallel R jobs
    • Loading Julia, packages, and running it in a job
  • Containers

    • Portable, reproducible containers for HPC
  • MPI implementations

    • The MPI library Anunna is built around
    • Intel MPI, with the Intel compilers and MKL

Graphical Interface Applications