<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.anunna.wur.nl/index.php?action=history&amp;feed=atom&amp;title=IntelMPI</id>
	<title>IntelMPI - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.anunna.wur.nl/index.php?action=history&amp;feed=atom&amp;title=IntelMPI"/>
	<link rel="alternate" type="text/html" href="https://wiki.anunna.wur.nl/index.php?title=IntelMPI&amp;action=history"/>
	<updated>2026-08-09T12:12:39Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.45.3</generator>
	<entry>
		<id>https://wiki.anunna.wur.nl/index.php?title=IntelMPI&amp;diff=3035&amp;oldid=prev</id>
		<title>Honfi001 at 14:46, 5 August 2026</title>
		<link rel="alternate" type="text/html" href="https://wiki.anunna.wur.nl/index.php?title=IntelMPI&amp;diff=3035&amp;oldid=prev"/>
		<updated>2026-08-05T14:46:12Z</updated>

		<summary type="html">&lt;p&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;en&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;← Older revision&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;Revision as of 14:46, 5 August 2026&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l44&quot;&gt;Line 44:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 44:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;module list           # everything loaded, dependencies included&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;module list           # everything loaded, dependencies included&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;mpirun -version       # which Intel MPI is on your PATH&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;mpirun -version       # which Intel MPI is on your PATH&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;&amp;lt;/syntaxhighlight&amp;gt;&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;If your software has its own module, load &#039;&#039;&#039;that&#039;&#039;&#039; instead and let it pull the matching MPI in as a dependency. A package built with this toolchain has a name ending in &amp;lt;code&amp;gt;-intel-2024a&amp;lt;/code&amp;gt; or similar, which tells you what you are getting:&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;&amp;lt;syntaxhighlight lang=&quot;bash&quot;&amp;gt;&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;module load 2024&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;module load &amp;lt;YourSoftware&amp;gt;/&amp;lt;version&amp;gt;    # Intel MPI arrives as a dependency&lt;/del&gt;&lt;/div&gt;&lt;/td&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-added&quot;&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;lt;/syntaxhighlight&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;lt;/syntaxhighlight&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l138&quot;&gt;Line 138:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 131:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;-genv&amp;lt;/code&amp;gt; is still the habit worth keeping&amp;#039;&amp;#039;&amp;#039; for anything the run depends on, as in the script above. It does not rely on which layer happens to be propagating, it states the intent in the launch line, it survives a site or script that restricts propagation later, and it is the first thing to try if a variable does mysteriously fail to arrive.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;-genv&amp;lt;/code&amp;gt; is still the habit worth keeping&amp;#039;&amp;#039;&amp;#039; for anything the run depends on, as in the script above. It does not rely on which layer happens to be propagating, it states the intent in the launch line, it survives a site or script that restricts propagation later, and it is the first thing to try if a variable does mysteriously fail to arrive.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Intel MPI, MKL, and &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;our &lt;/del&gt;AMD nodes ==&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;== Intel MPI, MKL, and AMD nodes ==&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;Most of Anunna&amp;#039;s compute is AMD, and the Intel stack has a well-known wrinkle there. It is worth being precise about where it lives: &amp;#039;&amp;#039;&amp;#039;the issue is in MKL, not in Intel MPI.&amp;#039;&amp;#039;&amp;#039; But MKL arrives with the same &amp;lt;code&amp;gt;intel/&amp;lt;year&amp;gt;a&amp;lt;/code&amp;gt; module, and &amp;quot;our Intel build is slow on the AMD nodes&amp;quot; is usually this, so it belongs on this page.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;Most of Anunna&amp;#039;s compute is AMD, and the Intel stack has a well-known wrinkle there. It is worth being precise about where it lives: &amp;#039;&amp;#039;&amp;#039;the issue is in MKL, not in Intel MPI.&amp;#039;&amp;#039;&amp;#039; But MKL arrives with the same &amp;lt;code&amp;gt;intel/&amp;lt;year&amp;gt;a&amp;lt;/code&amp;gt; module, and &amp;quot;our Intel build is slow on the AMD nodes&amp;quot; is usually this, so it belongs on this page.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l151&quot;&gt;Line 151:&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;Line 144:&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;lt;/syntaxhighlight&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;lt;/syntaxhighlight&amp;gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;Before assuming this is your problem, two things are worth checking, because often it is not. The shim only &#039;&#039;widens arithmetic&#039;&#039;, so it pays off when MKL is doing &#039;&#039;&#039;compute-bound &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Level-3 &lt;/del&gt;work&#039;&#039;&#039; &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;such as &amp;lt;code&amp;gt;DGEMM&amp;lt;/code&amp;gt;&lt;/del&gt;. It does nothing when MKL is only handling bandwidth-bound &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;Level-1/2 vector kernels&lt;/del&gt;, and nothing at all if your library does its own arithmetic &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;— &lt;/del&gt;PETSc&lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;, for instance, performs its sparse matrix–vector products itself, so MKL never sees the dominant cost of a Krylov solve&lt;/del&gt;.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;Before assuming this is your problem, two things are worth checking, because often it is not. The shim only &#039;&#039;widens arithmetic&#039;&#039;, so it pays off when MKL is doing &#039;&#039;&#039;compute-bound work&#039;&#039;&#039;. It does nothing when MKL is only handling bandwidth-bound &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;work&lt;/ins&gt;, and nothing at all if your library does its own arithmetic &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;e.g. &lt;/ins&gt;PETSc.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;br&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;lt;code&amp;gt;MKL_VERBOSE=1&amp;lt;/code&amp;gt; is the quick way to find out: it prints every MKL call, so you can see whether MKL is on your hot path at all and what it is being asked to do. Keep it off any run you intend to time — at high call counts the logging is itself a real cost.&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot;&gt;&lt;/td&gt;&lt;td style=&quot;background-color: #f8f9fa; color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #eaecf0; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&amp;lt;code&amp;gt;MKL_VERBOSE=1&amp;lt;/code&amp;gt; is the quick way to find out: it prints every MKL call, so you can see whether MKL is on your hot path at all and what it is being asked to do. Keep it off any run you intend to time — at high call counts the logging is itself a real cost.&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;/table&gt;</summary>
		<author><name>Honfi001</name></author>
	</entry>
	<entry>
		<id>https://wiki.anunna.wur.nl/index.php?title=IntelMPI&amp;diff=3034&amp;oldid=prev</id>
		<title>Honfi001: Created page with &quot;&#039;&#039;&#039;Intel MPI&#039;&#039;&#039; is the other MPI implementation available on Anunna, alongside OpenMPI. It arrives as part of the Intel toolchain, together with the Intel compilers and the Intel Math Kernel Library (MKL).  Everything general about MPI — that it runs one program as many cooperating processes, that it has to be built into the program, and when it is the right tool at all — is covered on Multi-Process Workflows. This page is the Intel-sp...&quot;</title>
		<link rel="alternate" type="text/html" href="https://wiki.anunna.wur.nl/index.php?title=IntelMPI&amp;diff=3034&amp;oldid=prev"/>
		<updated>2026-08-05T14:40:48Z</updated>

		<summary type="html">&lt;p&gt;Created page with &amp;quot;&amp;#039;&amp;#039;&amp;#039;Intel MPI&amp;#039;&amp;#039;&amp;#039; is the other MPI implementation available on Anunna, alongside &lt;a href=&quot;/OpenMPI&quot; title=&quot;OpenMPI&quot;&gt;OpenMPI&lt;/a&gt;. It arrives as part of the Intel toolchain, together with the Intel compilers and the Intel Math Kernel Library (MKL).  Everything general about MPI — that it runs one program as many cooperating processes, that it has to be built into the program, and when it is the right tool at all — is covered on &lt;a href=&quot;/Workflows/Multi-Process&quot; title=&quot;Workflows/Multi-Process&quot;&gt;Multi-Process Workflows&lt;/a&gt;. This page is the Intel-sp...&amp;quot;&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;&amp;#039;&amp;#039;&amp;#039;Intel MPI&amp;#039;&amp;#039;&amp;#039; is the other MPI implementation available on Anunna, alongside [[OpenMPI]]. It arrives as part of the Intel toolchain, together with the Intel compilers and the Intel Math Kernel Library (MKL).&lt;br /&gt;
&lt;br /&gt;
Everything general about MPI — that it runs one program as many cooperating processes, that it has to be built into the program, and when it is the right tool at all — is covered on [[Workflows/Multi-Process|Multi-Process Workflows]]. This page is the Intel-specific practice: which modules exist, how to launch, how the pinning controls differ from OpenMPI&amp;#039;s, and one thing that catches people out on our AMD nodes.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;Which should you use?&amp;#039;&amp;#039;&amp;#039; Whichever your program was built with. That is not a dodge — mixing MPI implementations between build and run does not work, so in practice the software chooses for you. Where you genuinely have a choice, Intel MPI and MKL are the natural pairing for code that leans on Intel&amp;#039;s maths libraries, and [[OpenMPI]] is the default for everything else here.&lt;br /&gt;
&lt;br /&gt;
== What is available on Anunna ==&lt;br /&gt;
&lt;br /&gt;
Software is built with [https://easybuild.io EasyBuild] and grouped into [[Environment Modules|buckets]]; a bucket has to be loaded before you can load anything from it. The Intel toolchain follows the same yearly generations as &amp;lt;code&amp;gt;foss&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Bucket !! Toolchain !! Contains&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;2023&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;intel/2023a&amp;lt;/code&amp;gt; || Intel compilers, Intel MPI, MKL&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;2024&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;intel/2024a&amp;lt;/code&amp;gt; || Intel compilers, Intel MPI, MKL&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;2025&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;intel/2025a&amp;lt;/code&amp;gt; || Intel compilers, Intel MPI, MKL&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;intel/&amp;lt;year&amp;gt;a&amp;lt;/code&amp;gt; is the counterpart of &amp;lt;code&amp;gt;foss/&amp;lt;year&amp;gt;a&amp;lt;/code&amp;gt;: one module that brings the whole stack — compilers, MPI, MKL and FFTW. There are leaner ways in, mirroring &amp;lt;code&amp;gt;gompi&amp;lt;/code&amp;gt; on the &amp;lt;code&amp;gt;foss&amp;lt;/code&amp;gt; side:&lt;br /&gt;
&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;iimpi/&amp;lt;year&amp;gt;a&amp;lt;/code&amp;gt;&amp;#039;&amp;#039;&amp;#039; — Intel compilers and Intel MPI, without MKL. The lean choice when you only need MPI.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;impi/&amp;lt;version&amp;gt;&amp;lt;/code&amp;gt;&amp;#039;&amp;#039;&amp;#039; — the MPI library alone. In the 2025 bucket this is &amp;lt;code&amp;gt;impi/2021.15.0-intel-compilers-2025.1.1&amp;lt;/code&amp;gt;; note that Intel MPI&amp;#039;s own version numbering (2021.x) does not track the toolchain year.&lt;br /&gt;
&lt;br /&gt;
To see what a bucket actually offers, ask the module system — &amp;lt;code&amp;gt;module key&amp;lt;/code&amp;gt; reads Lmod&amp;#039;s cache, so it needs no bucket loaded and searches every bucket at once:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
module key intel      # the toolchains&lt;br /&gt;
module key impi       # the MPI library on its own&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Loading it ===&lt;br /&gt;
&lt;br /&gt;
Bucket first, then the module — the same two steps as any other software here:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
module load 2024 intel/2024a&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Then confirm what you got, which is worth the habit:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
module list           # everything loaded, dependencies included&lt;br /&gt;
mpirun -version       # which Intel MPI is on your PATH&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If your software has its own module, load &amp;#039;&amp;#039;&amp;#039;that&amp;#039;&amp;#039;&amp;#039; instead and let it pull the matching MPI in as a dependency. A package built with this toolchain has a name ending in &amp;lt;code&amp;gt;-intel-2024a&amp;lt;/code&amp;gt; or similar, which tells you what you are getting:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
module load 2024&lt;br /&gt;
module load &amp;lt;YourSoftware&amp;gt;/&amp;lt;version&amp;gt;    # Intel MPI arrives as a dependency&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Keep a job inside one bucket. Do not load &amp;lt;code&amp;gt;intel&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;foss&amp;lt;/code&amp;gt; in the same job — they provide competing compilers, MPI libraries and BLAS, and the result is unpredictable rather than merely slow. If you want a clean starting point, &amp;lt;code&amp;gt;module reset&amp;lt;/code&amp;gt; returns you to the system defaults; the &amp;lt;code&amp;gt;slurm&amp;lt;/code&amp;gt; module is &amp;#039;&amp;#039;sticky&amp;#039;&amp;#039; and survives it either way.&lt;br /&gt;
&lt;br /&gt;
== Running on a single node ==&lt;br /&gt;
&lt;br /&gt;
Nothing runs on the login nodes. Put the work in a script, keep data on Lustre under &amp;lt;code&amp;gt;$myScratch&amp;lt;/code&amp;gt;, and submit with &amp;lt;code&amp;gt;sbatch&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash -l&lt;br /&gt;
#SBATCH --job-name=impi_single&lt;br /&gt;
#SBATCH --nodes=1&lt;br /&gt;
#SBATCH --ntasks=64&lt;br /&gt;
#SBATCH --cpus-per-task=1&lt;br /&gt;
#SBATCH --mem-per-cpu=2G        # per CORE, not per node&lt;br /&gt;
#SBATCH --time=01:00:00&lt;br /&gt;
#SBATCH --output=%x-%j.out&lt;br /&gt;
&lt;br /&gt;
cd &amp;quot;$myScratch/my_simulation&amp;quot;&lt;br /&gt;
&lt;br /&gt;
module load 2024 intel/2024a&lt;br /&gt;
&lt;br /&gt;
mpirun -np ${SLURM_NTASKS} ./my_mpi_program --input data.nc&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Two things carried over from [[OpenMPI]], for the same reasons: pass &amp;lt;code&amp;gt;${SLURM_NTASKS}&amp;lt;/code&amp;gt; rather than hard-coding the rank count, and use &amp;lt;code&amp;gt;--mem-per-cpu&amp;lt;/code&amp;gt; rather than &amp;lt;code&amp;gt;--mem&amp;lt;/code&amp;gt;, which is a per-&amp;#039;&amp;#039;node&amp;#039;&amp;#039; request divided among all the ranks on that node.&lt;br /&gt;
&lt;br /&gt;
Note the &amp;lt;code&amp;gt;#!/bin/bash -l&amp;lt;/code&amp;gt;. A login shell is what initialises Lmod, so without it &amp;lt;code&amp;gt;module load&amp;lt;/code&amp;gt; can fail quietly. For the same reason, never launch an MPI job with &amp;lt;code&amp;gt;sbatch --wrap&amp;lt;/code&amp;gt; — that runs under &amp;lt;code&amp;gt;dash&amp;lt;/code&amp;gt;, where &amp;lt;code&amp;gt;module&amp;lt;/code&amp;gt; does not exist.&lt;br /&gt;
&lt;br /&gt;
On one node the ranks talk through shared memory and there is nothing to configure.&lt;br /&gt;
&lt;br /&gt;
== Running across several nodes ==&lt;br /&gt;
&lt;br /&gt;
Anunna&amp;#039;s fabric is &amp;#039;&amp;#039;&amp;#039;Omni-Path (OPA100)&amp;#039;&amp;#039;&amp;#039;, and both MPI implementations reach it through &amp;#039;&amp;#039;&amp;#039;libfabric&amp;#039;&amp;#039;&amp;#039;. The provider that works reliably here is &amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;psm2&amp;lt;/code&amp;gt;&amp;#039;&amp;#039;&amp;#039;; the &amp;lt;code&amp;gt;opx&amp;lt;/code&amp;gt; provider is broken for inter-node traffic on this hardware. So an inter-node job should name the provider rather than trust auto-selection:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
#!/bin/bash -l&lt;br /&gt;
#SBATCH --job-name=impi_multi&lt;br /&gt;
#SBATCH --nodes=4&lt;br /&gt;
#SBATCH --ntasks-per-node=64&lt;br /&gt;
#SBATCH --cpus-per-task=1&lt;br /&gt;
#SBATCH --mem-per-cpu=2G&lt;br /&gt;
#SBATCH --time=02:00:00&lt;br /&gt;
#SBATCH --output=%x-%j.out&lt;br /&gt;
&lt;br /&gt;
cd &amp;quot;$myScratch/my_simulation&amp;quot;&lt;br /&gt;
&lt;br /&gt;
module load 2024 intel/2024a&lt;br /&gt;
&lt;br /&gt;
mpirun -genv I_MPI_OFI_PROVIDER psm2 -genv FI_PROVIDER psm2 \&lt;br /&gt;
       -np ${SLURM_NTASKS} ./my_mpi_program --input data.nc&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To see what was actually selected — provider, pinning, rank layout — turn up Intel MPI&amp;#039;s own diagnostics:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
export I_MPI_DEBUG=4&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
That prints the chosen fabric and the rank-to-core map at start-up, and is the fastest way to confirm a setting took effect instead of assuming it did.&lt;br /&gt;
&lt;br /&gt;
=== Getting variables to the other nodes ===&lt;br /&gt;
&lt;br /&gt;
Ranks on remote nodes are started fresh, so any variable your program or the fabric depends on has to reach them. Intel MPI&amp;#039;s controls are:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Option !! Effect&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;-genv &amp;lt;VAR&amp;gt; &amp;lt;value&amp;gt;&amp;lt;/code&amp;gt; || Set one variable for &amp;#039;&amp;#039;&amp;#039;all&amp;#039;&amp;#039;&amp;#039; ranks. The explicit, always-safe form.&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;-genvall&amp;lt;/code&amp;gt; || Pass the whole launching environment to all ranks. This is the default.&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;-genvlist &amp;lt;a,b,c&amp;gt;&amp;lt;/code&amp;gt; || Pass only the named variables.&lt;br /&gt;
|-&lt;br /&gt;
| &amp;lt;code&amp;gt;-genvnone&amp;lt;/code&amp;gt; || Pass nothing.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In practice &amp;#039;&amp;#039;&amp;#039;a plain &amp;lt;code&amp;gt;export&amp;lt;/code&amp;gt; in your job script does reach every rank&amp;#039;&amp;#039;&amp;#039; — tested on Anunna across two nodes.&lt;br /&gt;
&lt;br /&gt;
It arrives whatever you do to stop it: restricting Intel MPI&amp;#039;s own propagation with &amp;lt;code&amp;gt;-genvlist PATH&amp;lt;/code&amp;gt;, which should have excluded the test variable, delivered it to both nodes anyway. [[OpenMPI]] 5.0.7 behaved identically — with &amp;lt;code&amp;gt;-x&amp;lt;/code&amp;gt;, without it, and even with Slurm&amp;#039;s own environment forwarding suppressed.&lt;br /&gt;
&lt;br /&gt;
The practical reading is that &amp;#039;&amp;#039;&amp;#039;inside a Slurm allocation both implementations get your environment to the ranks&amp;#039;&amp;#039;&amp;#039;, by way of the process-management layer (PMIx and the launcher&amp;#039;s job description) rather than the batch environment. On the OpenMPI side that was pinned down: removing the variable from &amp;lt;code&amp;gt;mpirun&amp;lt;/code&amp;gt;&amp;#039;s own environment is the only thing that stopped it arriving — see [[OpenMPI#Exporting variables to the other nodes: the -x flag|the OpenMPI page]] for the full result.&lt;br /&gt;
&lt;br /&gt;
None of which is worth depending on in a job script. It held for the current toolchains and not necessarily for older ones, so say what your run needs explicitly with &amp;lt;code&amp;gt;-genv&amp;lt;/code&amp;gt; and the question stops mattering.&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;-genv&amp;lt;/code&amp;gt; is still the habit worth keeping&amp;#039;&amp;#039;&amp;#039; for anything the run depends on, as in the script above. It does not rely on which layer happens to be propagating, it states the intent in the launch line, it survives a site or script that restricts propagation later, and it is the first thing to try if a variable does mysteriously fail to arrive.&lt;br /&gt;
&lt;br /&gt;
== Intel MPI, MKL, and our AMD nodes ==&lt;br /&gt;
&lt;br /&gt;
Most of Anunna&amp;#039;s compute is AMD, and the Intel stack has a well-known wrinkle there. It is worth being precise about where it lives: &amp;#039;&amp;#039;&amp;#039;the issue is in MKL, not in Intel MPI.&amp;#039;&amp;#039;&amp;#039; But MKL arrives with the same &amp;lt;code&amp;gt;intel/&amp;lt;year&amp;gt;a&amp;lt;/code&amp;gt; module, and &amp;quot;our Intel build is slow on the AMD nodes&amp;quot; is usually this, so it belongs on this page.&lt;br /&gt;
&lt;br /&gt;
MKL checks the &amp;#039;&amp;#039;&amp;#039;CPU vendor&amp;#039;&amp;#039;&amp;#039; at run time. On a non-Intel CPU, kernels without specific Zen coverage can fall back to an &amp;#039;&amp;#039;&amp;#039;SSE&amp;#039;&amp;#039;&amp;#039; code path instead of using &amp;#039;&amp;#039;&amp;#039;AVX2&amp;#039;&amp;#039;&amp;#039;, leaving much of the vector width unused. This is still the behaviour in MKL 2025.x — it has not been quietly fixed. The old &amp;lt;code&amp;gt;MKL_DEBUG_CPU_TYPE&amp;lt;/code&amp;gt; workaround is gone, removed back in MKL 2020 Update 1, so anything you read recommending it is out of date.&lt;br /&gt;
&lt;br /&gt;
The current mitigation is a small shim, &amp;#039;&amp;#039;&amp;#039;&amp;lt;code&amp;gt;libfakeintel.so&amp;lt;/code&amp;gt;&amp;#039;&amp;#039;&amp;#039;, &amp;lt;code&amp;gt;LD_PRELOAD&amp;lt;/code&amp;gt;ed ahead of MKL. It overrides the vendor test (&amp;lt;code&amp;gt;mkl_serv_intel_cpu_true&amp;lt;/code&amp;gt;) so that it returns true, and MKL then dispatches its AVX2 kernels:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
export LD_PRELOAD=/path/to/libfakeintel.so&lt;br /&gt;
mpirun -np ${SLURM_NTASKS} ./my_mpi_program&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Before assuming this is your problem, two things are worth checking, because often it is not. The shim only &amp;#039;&amp;#039;widens arithmetic&amp;#039;&amp;#039;, so it pays off when MKL is doing &amp;#039;&amp;#039;&amp;#039;compute-bound Level-3 work&amp;#039;&amp;#039;&amp;#039; such as &amp;lt;code&amp;gt;DGEMM&amp;lt;/code&amp;gt;. It does nothing when MKL is only handling bandwidth-bound Level-1/2 vector kernels, and nothing at all if your library does its own arithmetic — PETSc, for instance, performs its sparse matrix–vector products itself, so MKL never sees the dominant cost of a Krylov solve.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;MKL_VERBOSE=1&amp;lt;/code&amp;gt; is the quick way to find out: it prints every MKL call, so you can see whether MKL is on your hot path at all and what it is being asked to do. Keep it off any run you intend to time — at high call counts the logging is itself a real cost.&lt;br /&gt;
&lt;br /&gt;
If MKL turns out to suit your workload poorly on the AMD nodes, the &amp;lt;code&amp;gt;foss&amp;lt;/code&amp;gt; toolchain&amp;#039;s OpenBLAS-based BLAS is worth benchmarking against.&lt;br /&gt;
&lt;br /&gt;
== Advanced: pinning and rank placement ==&lt;br /&gt;
&lt;br /&gt;
Intel MPI and OpenMPI express placement differently: OpenMPI takes command-line flags, Intel MPI reads &amp;lt;code&amp;gt;I_MPI_*&amp;lt;/code&amp;gt; environment variables. The intents map across cleanly, so if you know one dialect this table gives you the other:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Intent !! Intel MPI !! OpenMPI&lt;br /&gt;
|-&lt;br /&gt;
| One rank per physical core, no migration || &amp;lt;code&amp;gt;I_MPI_PIN=1&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;I_MPI_PIN_DOMAIN=core&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;--bind-to core&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| Fill the node, neighbours adjacent || &amp;lt;code&amp;gt;I_MPI_PIN_ORDER=compact&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;--map-by core&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| Under-subscribed, spread across all NUMA domains || &amp;lt;code&amp;gt;I_MPI_PIN_ORDER=scatter&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;--map-by numa --bind-to core&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| One rank per chiplet / L3 slice || &amp;lt;code&amp;gt;I_MPI_PIN_DOMAIN=cache3&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;--map-by l3cache --bind-to core&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| Print the map actually used || &amp;lt;code&amp;gt;I_MPI_DEBUG=4&amp;lt;/code&amp;gt; || &amp;lt;code&amp;gt;--report-bindings&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Three rules that matter more than the individual settings:&lt;br /&gt;
&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Verify, do not assume.&amp;#039;&amp;#039;&amp;#039; Print the map with &amp;lt;code&amp;gt;I_MPI_DEBUG=4&amp;lt;/code&amp;gt; before and after any change. Placement bugs do not announce themselves; the job runs and is simply slower than it should be.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;One source of truth.&amp;#039;&amp;#039;&amp;#039; Pin with the MPI launcher &amp;#039;&amp;#039;&amp;#039;or&amp;#039;&amp;#039;&amp;#039; with Slurm (&amp;lt;code&amp;gt;--cpu-bind&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;--distribution&amp;lt;/code&amp;gt;), never both. Two pinners fighting each other produce a nonsense map.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Spread suits memory-bound work.&amp;#039;&amp;#039;&amp;#039; AMD sockets are built from 8-core chiplets, each with its own cache slice and share of the memory channels. A bandwidth-bound job saturates memory well before every core is busy, so spreading a reduced number of ranks across all NUMA domains keeps every memory channel active, while packing them together leaves most idle. Asking for ranks in multiples of 8 keeps chiplets evenly filled.&lt;br /&gt;
&lt;br /&gt;
For the reasoning behind that last point, and the equivalent OpenMPI syntax, see [[OpenMPI#Advanced: pinning and rank placement|the OpenMPI page]].&lt;br /&gt;
&lt;br /&gt;
== See also ==&lt;br /&gt;
&lt;br /&gt;
* [[OpenMPI]] — the other MPI here, and the default for most software.&lt;br /&gt;
* [[Workflows/Multi-Process|Multi-Process Workflows]] — whether MPI is the right shape for your work.&lt;br /&gt;
* [[Environment Modules]] — buckets, and finding the module you need.&lt;br /&gt;
* [[Compute Hardware Overview]] — what the nodes actually are.&lt;br /&gt;
* [[Batch Jobs]] — writing and submitting job scripts.&lt;br /&gt;
* [[Scheduler Overview (Slurm)]] — how the scheduler allocates what you ask for.&lt;/div&gt;</summary>
		<author><name>Honfi001</name></author>
	</entry>
</feed>