mirror of
https://github.com/official-stockfish/Stockfish.git
synced 2026-07-22 20:57:10 +00:00
This enables different Stockfish processes that use the same weights to use the same memory. The approach establishes equivalence by memory content, and is compatible with NUMA replication. The benefit of sharing is reduced memory usage and a speedup thanks to improved (inter-process) caching of the network in the CPUs cache, and thus reduced bandwidth usage to main memory. Even though this change doesn't benefit a user running a single process, this helps on fishtest or e.g. for Lichess, when multiple games run concurrently, or multiple positions are analyzed in parallel. This concept was probably first introduced in the Monty engine (https://github.com/official-monty/Monty/pull/62), after a discussion in https://github.com/official-stockfish/fishtest/issues/2077 on the issue of memory pressure. Measurements based on Torch (https://github.com/user-attachments/files/21386224/verbatim.pdf) further suggested that large gains were possible. Multiple other engines have adopted this 'verbatim' format as well. The implementation here adds the flexibility needed for SF, for example, retains the ability to bundle compressed networks with the binary, to load nets by uci option, and to distribute the shared nets to the proper NUMA region. This flexibility comes with a fair amount of complexity in the implementation, such as OS specific code, and fallback code. For most users this should be transparent. However, for example, those running docker containers should ensure the `--ipc` flag is set correctly, and `--shm-size` is sufficiently large. The benefits of this patch significantly depend on hardware, with systems with many cores and a large (O(150MB), the net size) L3 cache benefitting typically most. On such systems SF speedups (as measured via nps playing games with large concurrency but just 1 thread) can be 38%, which results in master vs. patch Elo which gains about 25 Elo. ``` # PLAYER : RATING ERROR POINTS PLAYED (%) 1 shared_memoryPR : 24.8 1.9 39432.0 73728 53 2 master : 0.0 ---- 34296.0 73728 47 ``` In a multithreaded setup, where weights are already shared, that benefit is smaller, for example on the same HW as above, but with 8t for each side. ``` # PLAYER : RATING ERROR POINTS PLAYED (%) 1 shared_memoryPR : 5.2 3.5 9351.0 18432 51 2 master : 0.0 ---- 9081.0 18432 49 ``` On fishtest with a typical hardware mix of our contributors, the following was measured: STC, 60k games https://tests.stockfishchess.org/tests/view/69074a49ea4b268f1fac236c Elo: 4.69 ± 1.4 (95%) LOS: 100.0% Total: 60000 W: 16085 L: 15275 D: 28640 Ptnml(0-2): 154, 6440, 16053, 7148, 205 nElo: 9.38 ± 2.8 (95%) PairsRatio: 1.12 To verify correctness with a single process on a NUMA architecture, speedtest was used, confirming near equivalence: ``` master: Average (over 10): 296236186 shared_memory: Average (over 10): 295769332 ``` Currently, using large pages for the shared network weights is not always possible, which can lead to a small slowdown (1-2%), in case a single process is run. closes https://github.com/official-stockfish/Stockfish/pull/6173 No functional change Co-authored-by: disservin <disservin.social@gmail.com> Co-authored-by: Joost VandeVondele <Joost.VandeVondele@gmail.com>
200 lines
5.3 KiB
C++
200 lines
5.3 KiB
C++
/*
|
|
Stockfish, a UCI chess playing engine derived from Glaurung 2.1
|
|
Copyright (C) 2004-2025 The Stockfish developers (see AUTHORS file)
|
|
|
|
Stockfish is free software: you can redistribute it and/or modify
|
|
it under the terms of the GNU General Public License as published by
|
|
the Free Software Foundation, either version 3 of the License, or
|
|
(at your option) any later version.
|
|
|
|
Stockfish is distributed in the hope that it will be useful,
|
|
but WITHOUT ANY WARRANTY; without even the implied warranty of
|
|
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
|
GNU General Public License for more details.
|
|
|
|
You should have received a copy of the GNU General Public License
|
|
along with this program. If not, see <http://www.gnu.org/licenses/>.
|
|
*/
|
|
|
|
#include "memory.h"
|
|
|
|
#include <cstdlib>
|
|
|
|
#if __has_include("features.h")
|
|
#include <features.h>
|
|
#endif
|
|
|
|
#if defined(__linux__) && !defined(__ANDROID__)
|
|
#include <sys/mman.h>
|
|
#endif
|
|
|
|
#if defined(__APPLE__) || defined(__ANDROID__) || defined(__OpenBSD__) \
|
|
|| (defined(__GLIBCXX__) && !defined(_GLIBCXX_HAVE_ALIGNED_ALLOC) && !defined(_WIN32)) \
|
|
|| defined(__e2k__)
|
|
#define POSIXALIGNEDALLOC
|
|
#include <stdlib.h>
|
|
#endif
|
|
|
|
#ifdef _WIN32
|
|
#if _WIN32_WINNT < 0x0601
|
|
#undef _WIN32_WINNT
|
|
#define _WIN32_WINNT 0x0601 // Force to include needed API prototypes
|
|
#endif
|
|
|
|
#ifndef NOMINMAX
|
|
#define NOMINMAX
|
|
#endif
|
|
|
|
#include <ios> // std::hex, std::dec
|
|
#include <iostream> // std::cerr
|
|
#include <ostream> // std::endl
|
|
#include <windows.h>
|
|
|
|
// The needed Windows API for processor groups could be missed from old Windows
|
|
// versions, so instead of calling them directly (forcing the linker to resolve
|
|
// the calls at compile time), try to load them at runtime. To do this we need
|
|
// first to define the corresponding function pointers.
|
|
|
|
#endif
|
|
|
|
|
|
namespace Stockfish {
|
|
|
|
// Wrappers for systems where the c++17 implementation does not guarantee the
|
|
// availability of aligned_alloc(). Memory allocated with std_aligned_alloc()
|
|
// must be freed with std_aligned_free().
|
|
|
|
void* std_aligned_alloc(size_t alignment, size_t size) {
|
|
#if defined(_ISOC11_SOURCE)
|
|
return aligned_alloc(alignment, size);
|
|
#elif defined(POSIXALIGNEDALLOC)
|
|
void* mem = nullptr;
|
|
posix_memalign(&mem, alignment, size);
|
|
return mem;
|
|
#elif defined(_WIN32) && !defined(_M_ARM) && !defined(_M_ARM64)
|
|
return _mm_malloc(size, alignment);
|
|
#elif defined(_WIN32)
|
|
return _aligned_malloc(size, alignment);
|
|
#else
|
|
return std::aligned_alloc(alignment, size);
|
|
#endif
|
|
}
|
|
|
|
void std_aligned_free(void* ptr) {
|
|
|
|
#if defined(POSIXALIGNEDALLOC)
|
|
free(ptr);
|
|
#elif defined(_WIN32) && !defined(_M_ARM) && !defined(_M_ARM64)
|
|
_mm_free(ptr);
|
|
#elif defined(_WIN32)
|
|
_aligned_free(ptr);
|
|
#else
|
|
free(ptr);
|
|
#endif
|
|
}
|
|
|
|
// aligned_large_pages_alloc() will return suitably aligned memory,
|
|
// if possible using large pages.
|
|
|
|
#if defined(_WIN32)
|
|
|
|
static void* aligned_large_pages_alloc_windows([[maybe_unused]] size_t allocSize) {
|
|
|
|
return windows_try_with_large_page_priviliges(
|
|
[&](size_t largePageSize) {
|
|
// Round up size to full pages and allocate
|
|
allocSize = (allocSize + largePageSize - 1) & ~size_t(largePageSize - 1);
|
|
return VirtualAlloc(nullptr, allocSize, MEM_RESERVE | MEM_COMMIT | MEM_LARGE_PAGES,
|
|
PAGE_READWRITE);
|
|
},
|
|
[]() { return (void*) nullptr; });
|
|
}
|
|
|
|
void* aligned_large_pages_alloc(size_t allocSize) {
|
|
|
|
// Try to allocate large pages
|
|
void* mem = aligned_large_pages_alloc_windows(allocSize);
|
|
|
|
// Fall back to regular, page-aligned, allocation if necessary
|
|
if (!mem)
|
|
mem = VirtualAlloc(nullptr, allocSize, MEM_RESERVE | MEM_COMMIT, PAGE_READWRITE);
|
|
|
|
return mem;
|
|
}
|
|
|
|
#else
|
|
|
|
void* aligned_large_pages_alloc(size_t allocSize) {
|
|
|
|
#if defined(__linux__)
|
|
constexpr size_t alignment = 2 * 1024 * 1024; // 2MB page size assumed
|
|
#else
|
|
constexpr size_t alignment = 4096; // small page size assumed
|
|
#endif
|
|
|
|
// Round up to multiples of alignment
|
|
size_t size = ((allocSize + alignment - 1) / alignment) * alignment;
|
|
void* mem = std_aligned_alloc(alignment, size);
|
|
#if defined(MADV_HUGEPAGE)
|
|
madvise(mem, size, MADV_HUGEPAGE);
|
|
#endif
|
|
return mem;
|
|
}
|
|
|
|
#endif
|
|
|
|
bool has_large_pages() {
|
|
|
|
#if defined(_WIN32)
|
|
|
|
constexpr size_t page_size = 2 * 1024 * 1024; // 2MB page size assumed
|
|
void* mem = aligned_large_pages_alloc_windows(page_size);
|
|
if (mem == nullptr)
|
|
{
|
|
return false;
|
|
}
|
|
else
|
|
{
|
|
aligned_large_pages_free(mem);
|
|
return true;
|
|
}
|
|
|
|
#elif defined(__linux__)
|
|
|
|
#if defined(MADV_HUGEPAGE)
|
|
return true;
|
|
#else
|
|
return false;
|
|
#endif
|
|
|
|
#else
|
|
|
|
return false;
|
|
|
|
#endif
|
|
}
|
|
|
|
|
|
// aligned_large_pages_free() will free the previously memory allocated
|
|
// by aligned_large_pages_alloc(). The effect is a nop if mem == nullptr.
|
|
|
|
#if defined(_WIN32)
|
|
|
|
void aligned_large_pages_free(void* mem) {
|
|
|
|
if (mem && !VirtualFree(mem, 0, MEM_RELEASE))
|
|
{
|
|
DWORD err = GetLastError();
|
|
std::cerr << "Failed to free large page memory. Error code: 0x" << std::hex << err
|
|
<< std::dec << std::endl;
|
|
exit(EXIT_FAILURE);
|
|
}
|
|
}
|
|
|
|
#else
|
|
|
|
void aligned_large_pages_free(void* mem) { std_aligned_free(mem); }
|
|
|
|
#endif
|
|
} // namespace Stockfish
|