This document defines a repeatable framework for evaluating storage and data-path performance across StratiSTOR, StratiSERV, and HyperSERV environments. It is intended for infrastructure engineers, architects, partners, performance teams, and operators who need to measure throughput, IOPS, latency, scalability, protocol efficiency, and application-oriented behavior using controlled and reproducible test methods.
The framework combines synthetic and workload-oriented tools so performance can be evaluated at multiple layers: local and distributed storage, virtual-machine I/O, network file services, block protocols, object storage, high-performance and parallel filesystems, and the network paths connecting clients to the platform.
Terminology: StratiSTOR refers to the SteelDome distributed storage and data-services layer, StratiSERV refers to the virtualization and compute layer, and HyperSERV refers to converged StratiSTOR and StratiSERV deployments. Lustre is referenced directly where parallel filesystem testing is applicable.
IMPORTANT: Performance results are meaningful only when the test topology, protection policy, client count, network configuration, CPU/NUMA placement, cache state, working-set size, tool version, and exact test parameters are recorded. Results from different test conditions should not be compared as if they were equivalent.
Storage performance is the result of an entire data path rather than a single device specification. CPU resources, NUMA locality, memory, storage media, protection policy, metadata behavior, network bandwidth, NIC queues, protocol overhead, client concurrency, caching, virtualization, and application access patterns can all influence observed performance.
For that reason, this framework evaluates performance from several perspectives rather than relying on a single benchmark. SteelDome Perfprofiler and FIO provide controlled block and file I/O measurements; IOR evaluates parallel and distributed filesystem behavior; Frametest represents high-bandwidth media workflows; S3 Warp measures object-storage behavior; and iperf3 establishes whether the network can support the storage throughput expected from the system.
The same framework can be used for dedicated StratiSTOR clusters, StratiSERV virtual machines consuming external or StratiSTOR-backed storage, HyperSERV hyperconverged deployments, tiered over-the-wire testing, Kubernetes data paths, Lustre-based HPC environments, and AI infrastructure where storage performance must be considered together with high-speed networking and accelerated compute.
A useful benchmark must be reproducible and must isolate the component or path being measured. Before recording performance, document the test objective and determine whether the test is intended to measure media performance, cluster behavior, protocol overhead, client scalability, VM performance, object performance, or complete application data-path performance.
| TOOL | PRIMARY PURPOSE | KEY METRICS | BEST FIT |
|---|---|---|---|
| Perfprofiler | Automated FIO-based profiling across block sizes and access patterns | IOPS, throughput, latency | Broad performance characterization and repeatable comparison |
| FIO | Flexible synthetic block and file I/O generation | IOPS, bandwidth, latency, percentiles | Block devices, filesystems, VM disks, protocol-mounted storage |
| IOR | Distributed and parallel filesystem testing through MPI | Aggregate throughput and per-process statistics | HPC, Lustre, distributed filesystems, multi-client scale testing |
| Frametest | Media-frame and sustained streaming workload simulation | Frame rate, throughput, frame completion time | Media, post-production, rendering, high-bandwidth file workloads |
| S3 Warp | S3-compatible object benchmarking | PUT/GET throughput, operations, latency, concurrency | Object storage, backup, data lake, cloud-native and AI datasets |
| iperf3 | Network baseline and path validation | Bandwidth, retransmits, parallel-stream behavior | Front-end and cluster-network capability validation before storage tests |
The location of the load generator determines what the benchmark measures. Clearly identify the test layer in every published result.
| Node-Local Test | Measures a local device or local filesystem path. Useful for verifying hardware capability but does not represent distributed or protocol performance. |
| Cluster-Native Test | Runs against a StratiSTOR cluster-mounted data path and includes distributed placement, protection, metadata, and cluster-network behavior. |
| Over-the-Wire Protocol Test | Uses external clients through NFS, SMB, iSCSI, NVMe-oF, S3, or another presented service. This includes the client network, protocol stack, service layer, and storage backend. |
| Virtual Machine Test | Runs from a StratiSERV or HyperSERV guest and measures the VM storage path, virtual hardware, host scheduling, network/storage backend, and any configured cache or acceleration layers. |
| Parallel / HPC Test | Uses multiple clients or MPI processes against Lustre or another parallel data path to evaluate aggregate throughput and scaling. |
| Recovery-Under-Load Test | Intentionally measures performance during device, node, or site recovery. Report separately from normal steady-state benchmark results. |
| COMPONENT | TESTING GUIDANCE |
|---|---|
| CPU |
|
| MEMORY |
|
| NETWORK |
|
| STORAGE |
|
Validate network capability before interpreting storage results. A storage benchmark cannot exceed the effective capacity of the network path carrying its traffic, and high-speed systems can also be limited by CPU, PCIe bandwidth, NIC queue configuration, NUMA placement, or the benchmark client itself.
Server — eight independent iperf3 listeners on TCP 5101-5108:
for port in {5101..5108}; do
iperf3 -s -p "$port" &
done
waitClient — eight concurrent flows:
for i in {1..8}; do
iperf3 -c 192.168.100.43 -p $((5100 + i)) -w 1M -Z &
done
waitFor very high-bandwidth adapters, increase the number of independent flows gradually. 100/200/400/800 Gb environments may require additional processes or streams, multiple CPU cores, and careful NUMA placement before the link can be saturated.
A healthy test path should approach the practical payload capability of the link. Approximately 90% of nominal line rate is a useful target when the host, PCIe path, CPU, NIC, and test configuration are capable of driving it, but protocol overhead and hardware limitations can lower the achievable rate. If results are unexpectedly low, inspect the following before storage testing:
Useful validation commands:
ethtool <interface>
ethtool -i <interface>
ethtool -l <interface>
ethtool -x <interface>
ip -s link show <interface>
lspci -vv -s <NIC_PCI_ADDRESS>
numactl --hardwareSteelDome Perfprofiler uses FIO as its workload-generation engine and automates broad sweeps across block sizes, sequential and random access patterns, queue depth, thread count, working-set size, direct I/O, and runtime. It is useful when a complete performance profile is more informative than a single peak number.
Best use: establish a performance surface for a device, filesystem, StratiSTOR cluster path, VM disk, or mounted network filesystem and retain CSV/JSON data for later comparison and visualization.
FIO provides precise control over block size, read/write mix, queue depth, process count, runtime, I/O engine, direct I/O, CPU affinity, and output format. It is the primary low-level tool in this framework for testing block devices, filesystems, VM disks, and protocol-mounted storage.
Best use: reproduce application-oriented I/O profiles or isolate individual dimensions such as 4 KiB random IOPS, large-block sequential throughput, mixed read/write behavior, or latency at a controlled queue depth.
Use io_uring on current Linux systems where appropriate; retain libaio when comparing with older data or when the target environment requires it.
IOR is an MPI-based parallel I/O benchmark used to measure aggregate filesystem performance across multiple processes and clients. It is particularly valuable for Lustre, distributed filesystems, research/HPC workloads, and cluster-scale throughput testing.
Best use: determine how aggregate throughput scales as clients and processes are added, and identify whether the bottleneck resides in clients, metadata, storage nodes, network bandwidth, or the filesystem data path.
Frametest generates large streaming file workloads intended to resemble media-frame processing. It reports frame rate, throughput, and completion timing, which makes it useful for post-production, rendering, content processing, and other workloads where sustained file delivery is more meaningful than small-block IOPS.
S3 Warp is a load-generation and benchmarking tool for S3-compatible object storage. It can exercise GET, PUT, DELETE, LIST, mixed workloads, object-size distributions, and multiple concurrent clients while reporting throughput and latency.
Best use: measure object-service scalability and client concurrency for backup, data-lake, cloud-native, archive, analytics, and AI dataset workflows.
Perfprofiler is included with StratiSYSTEM. The product role determines the base directory. Use the path present on the system:
# StratiSTOR / HyperSERV storage role
ln -s /stratistor/modules/perftools/perfprofiler/perfprofiler /usr/local/bin/perfprofiler
# StratiSERV compute role
ln -s /stratiserv/modules/perftools/perfprofiler/perfprofiler /usr/local/bin/perfprofiler
Perfprofiler depends on FIO. On non-StratiSYSTEM clients, install a compatible FIO package before using Perfprofiler.
IOR binaries are included with StratiSYSTEM. Install the MPI runtime/development components required by the current build on every test client, then make the included binaries available in the executable path.
dnf install -y openmpi-devel
ln -sf /usr/lib64/openmpi/bin/mpirun /usr/local/bin/mpirun
ln -sf /usr/lib64/openmpi/bin/orted /usr/local/bin/orted
# Example for StratiSTOR / HyperSERV storage role
cp /stratistor/modules/perftools/ior/{ior,mdtest,md-workbench} /usr/local/bin/
export OMPI_ALLOW_RUN_AS_ROOT=1
export OMPI_ALLOW_RUN_AS_ROOT_CONFIRM=1
NOTE: Running MPI as root is convenient for controlled benchmark environments but is not a general production best practice. Use a dedicated test account where operational policy requires it.
FIO is included with StratiSYSTEM and is normally available from /usr/bin/fio. On compatible external Linux clients:
dnf install -y fioFrametest is included with StratiSYSTEM. Example:
# StratiSTOR / HyperSERV storage role
ln -s /stratistor/modules/perftools/frametest/frametest /usr/local/bin/frametest
# StratiSERV compute role
ln -s /stratiserv/modules/perftools/frametest/frametest /usr/local/bin/frametestThe StratiSYSTEM software bundle includes the S3 Warp package used by this framework. Install the RPM from the applicable product-role directory. Use rpm -Uvh when installing or replacing the packaged version:
rpm -Uvh /stratistor/modules/perftools/s3warp/s3_warp_linux_x86_64.rpmNFS provides shared file access for Linux/Unix and other NFS-capable clients. NFSv4.x typically uses TCP 2049. When benchmarking, record the negotiated protocol version, mount options, client count, rsize/wsize, attribute-cache behavior, and whether the test is intended to represent production settings or an intentionally cache-minimized benchmark.
SMB provides shared file access for Windows, Linux, and other SMB-capable clients. Modern SMB uses TCP 445. Record dialect, signing/encryption state, SMB Multichannel status, RSS/RDMA capability, client NIC count, and server interface configuration because these can materially affect throughput.
iSCSI provides block storage over TCP/IP and normally uses TCP 3260. Performance depends on initiator queueing, multipath configuration, NIC and CPU capability, target layout, and the filesystem or database layered above the block device.
NVMe over Fabrics provides high-performance remote block access. NVMe/TCP commonly uses TCP 4420; RDMA-based transports require the corresponding RDMA fabric configuration. Record transport type, multipath configuration, queue count, controller settings, NIC/NUMA placement, and network loss characteristics.
S3-compatible access is used for object workloads. The service endpoint and port are deployment-specific. The S3 Warp example later in this guide uses TLS endpoint port 15443; this is an example from the documented test environment, not a universal StratiSTOR S3 port.
Lustre is appropriate for highly parallel HPC, AI, research, and throughput-oriented workloads. IOR is the preferred tool in this framework for aggregate Lustre testing. Record client count, MPI process count, stripe configuration, network transport, metadata configuration, and the relationship between client, metadata, and object-storage resources.
Advanced internal testing can be performed against StratiSTOR-native cluster paths when required for engineering or support analysis. These paths are implementation-specific and may change between releases; public and customer-facing benchmark procedures should normally use the supported NFS, SMB, iSCSI, NVMe-oF, S3, Lustre, or workload-facing interfaces unless SteelDome support directs otherwise.
Example production-oriented mount:
mount -t nfs4 -o rw,hard,vers=4.1,proto=tcp,rsize=1048576,wsize=1048576 \
<STRATISTOR_IP_OR_VIP>:/nfs01 /mnt/<LOCAL_MOUNT_DIR>
For cache-minimized benchmark runs, additional mount options may be used, but document them explicitly because options such as noac and lookupcache=none can significantly change metadata behavior and may not represent normal production usage.
mount -t cifs //<STRATISTOR_IP_OR_VIP>/<SHARE_NAME> /mnt/smbshare \
-o username=<USERNAME>,vers=3.1.1
Avoid placing passwords directly on the command line in shared environments. Use a protected credentials file when appropriate. On Windows clients, validate SMB Multichannel and RSS/RDMA state before interpreting high-bandwidth test results.
Discovery — TCP 3260:
iscsiadm -m discovery -t sendtargets -p <TARGET_IP>:3260
Login:
iscsiadm -m node -T <TARGET_IQN> -p <TARGET_IP>:3260 --login
After login, verify the discovered device and multipath state before creating a filesystem or running destructive benchmarks.
Discovery — TCP 4420:
nvme discover -t tcp -a <TARGET_IP> -s 4420
Connect:
nvme connect -t tcp -n <TARGET_NQN> -a <TARGET_IP> -s 4420S3 Warp connects directly to the configured StratiSTOR S3 endpoint and does not require a filesystem mount. Confirm the endpoint, TLS configuration, bucket, access key, secret key, and service port before testing. The example in this guide uses TCP 15443 because that was the endpoint configured in the reference test environment.
Use the Lustre mount information generated for the deployed filesystem. Generic example:
mkdir -p /mnt/lustre
mount -t lustre <MGS_NID>:/<FILESYSTEM_NAME> /mnt/lustre
Before IOR testing, verify the intended client count, network transport, mount state, and stripe policy.
TEST COMMAND:
perfprofiler [blocksizes] [patterns] [iodepth] [runtime] [threads] [startdelay] [directio] [numfiles] [savecsv] [size] [testlocation] [statfilelocation]
<blocksizes> Comma-separated list of block sizes (e.g., 4K,16K) or 'all' for all block sizes (4K-4M)
<patterns> Test patterns: 'random', 'sequential', or 'all' for all patterns
<iodepth> IO depth [default: 32]
<runtime> Runtime for each test in seconds [default: 60]
<threads> Number of threads [default: 4]
<startdelay> Start delay in seconds [default: 0]
<directio> Bypass cache (0 - no, 1 - yes) [default: 1]
<numfiles> Number of test files [default: 1]
<savecsv> Save output in CSV format (0 - no, 1 - yes) [default: 1]
<size> Size of the working set (e.g., 16G, 32G, 128G)
<testlocation> Test device or file (e.g., /dev/sdz or /tmp/testfile)
<statfilelocation> Location to save stats file (e.g., /tmp/stats_file.json)
EXAMPLE TEST RUN
./perfprofiler all all 64 10 8 2 1 8 1 2G /tmp/testfile /tmp/stats_file.json
SteelDome Storage Profiler Tool v1.3
Test executing with the following parameters:
BLOCK SIZE: all
PATTERN: all
TEST SIZE: 2G
IO DEPTH: 64
PATTERN RUNTIME: 10 sec(s)
THREADS: 8 threads
START DELAY: 2 sec(s)
DIRECT IO: YES (Bypass cache)
TEST FILES: 8 file(s)
TEST LOCATION: /tmp/[auto-generated]
STATS LOCATION: /tmp/stats_file.json
CSV OUTPUT: YES
Pattern Operation Block Size IOPS Throughput (MB/s) Latency (ms)
--------------------------------------------------------------------------------------
write-random write 4K 82012 320 6.197
write-random write 8K 76021 594 6.684
write-random write 16K 91273 1426 5.597
write-random write 32K 74856 2339 6.826
write-random write 64K 37675 2355 13.567
write-random write 128K 17072 2134 29.944
write-random write 256K 8140 2035 62.748
write-random write 512K 3523 1761 145.095
write-random write 1M 2380 2380 214.049
write-random write 2M 1241 2482 411.163
write-random write 4M 625 2500 808.705
read-random read 4K 188152 735 2.718
read-random read 8K 167145 1306 3.060
read-random read 16K 148142 2315 3.452
read-random read 32K 138574 4330 3.691
read-random read 64K 86128 5383 5.940
read-random read 128K 41811 5226 12.233
read-random read 256K 20993 5248 24.366
read-random read 512K 6458 3229 79.073
read-random read 1M 5465 5465 93.363
read-random read 2M 5326 10652 95.801
read-random read 4M 3247 12988 156.566
write-sequential write 4K 378626 1479 1.349
write-sequential write 8K 231467 1808 2.207
write-sequential write 16K 126857 1982 4.027
write-sequential write 32K 68329 2135 7.480
write-sequential write 64K 35010 2188 14.600
write-sequential write 128K 16376 2047 31.218
write-sequential write 256K 8188 2047 62.413
write-sequential write 512K 4265 2132 119.766
write-sequential write 1M 2163 2163 235.850
write-sequential write 2M 1035 2070 491.282
write-sequential write 4M 544 2174 912.029
read-sequential read 4K 443858 1734 1.152
read-sequential read 8K 371094 2899 1.377
read-sequential read 16K 287801 4497 1.776
read-sequential read 32K 250766 7836 2.039
read-sequential read 64K 183566 11473 2.786
read-sequential read 128K 114879 14360 4.452
read-sequential read 256K 71103 17776 7.193
read-sequential read 512K 40824 20412 12.523
read-sequential read 1M 21028 21028 24.298
read-sequential read 2M 8801 17603 57.883
read-sequential read 4M 3431 13723 146.434
The CSV output can be used to easily generate a 3D performance profile providing a graphical snapshot of the full spectrum of performance across all performance parameters.
IO PERFORMANCE PROFILE
THROUGHPUT PERFORMANCE PROFILE
WRITE TEST
TEST COMMAND:
frametest -w 4k -n10000 -t12 .
Explanation:
READ TEST
TEST COMMAND:
frametest -r 4k -n10000 -t12 .
Explanation:
NOTE: You must first run frametest write operations before running read operations as the read operation requires the frametest files to be in place.
EXAMPLE TEST RUN:
frametest -r 4k -n 5000 -t 20 .
Test parameters: -r -z49856 -n5000 -t20
Test duration: 18 secs
Frames transferred: 4918 (239445.125 MB)
Fastest frame: 52.669 ms (924.41 MB/s)
Slowest frame: 132.407 ms (367.71 MB/s)
Averaged details:
Open I/O Frame Data rate Frame rate
Last 1s: 0.454 ms 68.47 ms 3.64 ms 13369.12 MB/s 274.6 fps
5s: 0.436 ms 67.42 ms 3.69 ms 13191.83 MB/s 270.9 fps
30s: 0.382 ms 67.74 ms 3.67 ms 13276.03 MB/s 272.7 fps
Overall: 0.382 ms 67.74 ms 3.67 ms 13276.03 MB/s 272.7 fps
Histogram of frame completion times:
100% |
|
|
|
| *
| *
| *
| **
| **
| *******
+|----|-----|----|----|-----|----|----|-----|----|----|-----|----|
ms <0.1 .2 .5 1 2 5 10 20 50 100 200 500 >1s
Overall frame rate ..... 273.02 fps (13938416546 bytes/s)
Average file time ...... 68.095 ms
Shortest file time ..... 52.669 ms
Longest file time ...... 132.407 ms
Average open time ...... 0.380 ms
Shortest open time ..... 0.017 ms
Longest open time ...... 8.824 ms
Average read time ...... 67.7 ms
Shortest read time ..... 52.5 ms
Longest read time ...... 132.2 ms
Average close time ..... 0.006 ms
Shortest close time .... 0.001 ms
Longest close time ..... 0.027 ms
1. Small Sequential I/O Size Test (4 KiB)
TEST COMMAND:
mpirun -mca btl tcp -np 128 -H hyper01:16,hyper02:16,hyper03:16,hyper04:16,hyper05:16,hyper06:16,hyper07:16,hyper08:16 \
/stratistor/modules/perftools/ior/ior -t 4k -b 256m -s 64 -F -C -e -i 2 -o /stratistor/clustermounts/machines/ior-test/testfile
Explanation:
Executes the IOR tool across eight (8) test nodes with a total of 128 tasks, 16 clients per node, using 64 segments, with a transfer size of 4 KiB, block size of 256 MiB, and a total aggregate working set size of 2 TiB.
2. Medium Sequential I/O Size Test (32 KiB)
TEST COMMAND:
mpirun -mca btl tcp -np 128 -H hyper01:16,hyper02:16,hyper03:16,hyper04:16,hyper05:16,hyper06:16,hyper07:16,hyper08:16 \
/stratistor/modules/perftools/ior/ior -t 32k -b 256m -s 64 -F -C -e -i 2 -o /stratistor/clustermounts/machines/ior-test/testfile
Explanation:
Executes the IOR tool across eight (8) test nodes with a total of 128 tasks, 16 clients per node, using 64 segments, with a transfer size of 32 KiB, block size of 256 MiB, and a total aggregate working set size of 2 TiB.
3. Large Sequential I/O Size Test (1 MiB)
TEST COMMAND:
mpirun -mca btl tcp -np 128 -H hyper01:16,hyper02:16,hyper03:16,hyper04:16,hyper05:16,hyper06:16,hyper07:16,hyper08:16 \
/stratistor/modules/perftools/ior/ior -t 1m -b 256m -s 64 -F -C -e -i 2 -o /stratistor/clustermounts/machines/ior-test/testfile
Explanation:
Executes the IOR tool across eight (8) test nodes with a total of 128 tasks, 16 clients per node, using 64 segments, with a transfer size of 1 MiB, block size of 256 MiB, and a total aggregate working set size of 2 TiB.
4. Small Random I/O Size Test (4 KiB)
TEST COMMAND:
mpirun -mca btl tcp -np 128 -H hyper01:16,hyper02:16,hyper03:16,hyper04:16,hyper05:16,hyper06:16,hyper07:16,hyper08:16 \
/stratistor/modules/perftools/ior/ior -t 4k -b 256m -s 64 -F -C -e -i 2 -z -o /stratistor/clustermounts/machines/ior-test/testfile
Explanation:
Executes the IOR tool across eight (8) test nodes with a total of 128 tasks, 16 clients per node, using 64 segments, with a transfer size of 4 KiB, block size of 256 MiB, and a total aggregate working set size of 2 TiB.
5. Medium Random I/O Size Test (32 KiB)
TEST COMMAND:
mpirun -mca btl tcp -np 128 -H hyper01:16,hyper02:16,hyper03:16,hyper04:16,hyper05:16,hyper06:16,hyper07:16,hyper08:16 \
/stratistor/modules/perftools/ior/ior -t 32k -b 256m -s 64 -F -C -e -i 2 -z -o /stratistor/clustermounts/machines/ior-test/testfile
Explanation:
Executes the IOR tool across eight (8) test nodes with a total of 128 tasks, 16 clients per node, using 64 segments, with a transfer size of 32 KiB, block size of 256 MiB, and a total aggregate working set size of 2 TiB.
6. Large Random I/O Size Test (1 MiB)
TEST COMMAND:
mpirun -mca btl tcp -np 128 -H hyper01:16,hyper02:16,hyper03:16,hyper04:16,hyper05:16,hyper06:16,hyper07:16,hyper08:16 \
/stratistor/modules/perftools/ior/ior -t 1m -b 256m -s 64 -F -C -e -i 2 -z -o /stratistor/clustermounts/machines/ior-test/testfile
Explanation:
Executes the IOR tool across eight (8) test nodes with a total of 128 tasks, 16 clients per node, using 64 segments, with a transfer size of 1 MiB, block size of 256 MiB, and a total aggregate working set size of 2 TiB.
1. Small Sequential I/O Size Test (4 KiB)
TEST COMMAND:
fio --filename=/stratistor/clustermounts/machines/fio/fiotest --ioengine=libaio --time_based --group_reporting --refill_buffers --norandommap \
--randrepeat=0 --unlink=1 --name=write-random-write --nrfiles=1 --direct=1 --iodepth=64 --runtime=10 --numjobs=8 --rw=write --bs=4K --size=32G \
--startdelay=2 --output-format=json
Explanation:
Executes the FIO tool on a single host using 4 KiB block sizes, queue depth of 64 commands, cache disabled, 32 GiB working data set.
2. Small Random I/O Size Test (4 KiB)
TEST COMMAND:
fio --filename=/stratistor/clustermounts/machines/fio/fiotest --ioengine=libaio --time_based --group_reporting --refill_buffers --norandommap \
--randrepeat=0 --unlink=1 --name=write-random-write --nrfiles=1 --direct=1 --iodepth=64 --runtime=10 --numjobs=8 --rw=randwrite --bs=4K --size=32G \
--startdelay=2 --output-format=json
Explanation:
Executes the FIO tool on a single host using 4 KiB block sizes, queue depth of 64 commands, cache disabled, 32 GiB working data set.
CLIENT EXECUTION
Command on each test client node:
warp client
SERVER EXECUTION
Command on each test server node:
warp get --duration=3m --warp-client=10.184.25.73,10.184.27.23,10.184.19.151,10.184.18.81 --host=192.168.101.1:15443,192.168.101.2:15443,192.168.101.3:15443 \
--access-key=[S3 ACCESS KEY] --secret-key=[S3 SECRET KEY] --tls --bucket bucket-001 --insecure
IMPORTANT CONTEXT: The benchmark data below is preserved as a historical reference from the documented test environment. It was generated on the specific hardware, software release, topology, working set, and test parameters shown here. It should not be interpreted as a performance guarantee for StratiSYSTEM OS 4.9 or for different hardware and network configurations. Use the methodology in this guide to establish results for the actual deployment being evaluated.
The following reference results were generated on a high-performance HyperSERV hyperconverged deployment using the hardware listed below. They illustrate the behavior of that specific test environment under the documented workload and should be compared only with results produced using equivalent methodology, topology, client count, protection policy, cache state, and test parameters.
Each node in the cluster consists of a Supermicro SYS-221BT-HNC8R server equipped with:
2× Intel® Xeon® Gold 6438Y+ CPUs - 64 physical cores and 128 threads per node @2.0 GHz, offering a balanced mix of performance and power efficiency for multi-threaded workloads.
256 GB of DDR4 RAM, ensuring sufficient memory capacity to support virtualization, in-memory processing, and high-throughput tasks.
6× Samsung 3.84TB PCIe Gen5 NVMe SSDs (Model: MZWLO3T8HCLS), delivering ultra-fast storage performance with high IOPS, low latency, and sustained throughput, ideal for storage-intensive applications.
1× Mellanox ConnectX-7 400GbE network adapter (MT2910), providing extremely high network bandwidth and low-latency connectivity, enabling fast data movement across nodes and to external clients.
SteelDome StratiSYSTEM OS 3.3. These results are historical and are retained for reference; current StratiSYSTEM OS 4.9 deployments should be benchmarked independently using the current framework.
TEST COMMAND:
mpirun -mca btl tcp -np 256 -H hyper01:64,hyper02:64,hyper03:64,hyper04:64 /stratistor/modules/perftools/ior/ior \
-t 4k -b 256m -s 256 -F -C -e -u -o /stratistor/clustermounts/ior-test/testfile
IOR 4-NODE TEST RESULTS:
access bw(MiB/s) IOPS Latency(s) block(KiB) xfer(KiB)
------ --------- ---------- ---------- ---------- ---------
write 14,862 3,804,719 0.000067 256 4
read 39,161 10,029,039 0.000017 256 4
TEST COMMAND:
mpirun -mca btl tcp -np 256 -H hyper01:32,hyper02:32,hyper03:32,hyper04:32,hyper05:32,hyper06:32,hyper07:32,hyper08:32 /stratistor/modules/perftools/ior/ior \
-t 4k -b 256m -s 256 -F -C -e -u -o /stratistor/clustermounts/ior-test/testfile
IOR 8-NODE TEST RESULTS:
access bw(MiB/s) IOPS Latency(s) block(KiB) xfer(KiB)
------ --------- ---------- ---------- ---------- ---------
write 29,923 7,660,475 0.000017 256 4
read 94,769 24,262,194 0.000005 256 4
To evaluate system performance at scale, Frametest was executed concurrently across eight HyperSERV nodes, using 4K frames to reflect real-world demands in media and entertainment environments. This approach aligns with industry standards for testing high-throughput, low-latency infrastructure—ensuring the platform can reliably handle the intense concurrent data and graphics processing requirements typical of video rendering, post-production, and live content delivery pipelines.
WRITE TEST COMMANDS:
(nodes 1-8): frametest -w 4k -n5000 -t20 .
FRAMETEST WRITE THROUGHPUT TEST RESULTS:
NODE # Data Rate Framerate
------ ------------- ------------
1 5443.49 MB/s 111.8 fps
2 5525.70 MB/s 113.5 fps
3 5475.75 MB/s 112.5 fps
4 5559.42 MB/s 114.2 fps
5 5611.81 MB/s 115.3 fps
6 5492.07 MB/s 112.8 fps
7 5498.72 MB/s 112.9 fps
8 5603.56 MB/s 115.1 fps
-----------------------------------
TOTAL: 48,709.5 MB/s 1008.1 fps
READ TEST COMMANDS:
(nodes 1-8): frametest -r 4k -n5000 -t20 .
FRAMETEST READ THROUGHPUT TEST RESULTS:
NODE # Data Rate Framerate
------ ------------- -----------
1 13343.73 MB/s 274.1 fps
2 13290.64 MB/s 273.0 fps
3 13333.80 MB/s 273.9 fps
4 13343.54 MB/s 274.1 fps
5 13337.88 MB/s 273.9 fps
6 13336.59 MB/s 273.9 fps
7 13298.43 MB/s 273.2 fps
8 13441.37 MB/s 274.0 fps
----------------------------------
TOTAL: 106,727 MB/s 2190.1 fps
The following test results are based on a virtual machine deployed directly on the HyperSERV platform, utilizing the high-performance hardware configuration detailed above. This setup is designed to emulate real-world operating conditions, ensuring that the results reflect the true capabilities of the system in enterprise-grade environments.
FIO IOPS TEST COMMANDS:
fio --rw=read --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=4k --iodepth=128 --numjobs=12 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
fio --rw=randread --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=4k --iodepth=128 --numjobs=12 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
fio --rw=write --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=4k --iodepth=128 --numjobs=12 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
fio --rw=randwrite --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=4k --iodepth=128 --numjobs=12 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
FIO IOPS TEST RESULTS:
Test Type Operation Block Size IOPS Bandwidth
----------- ---------- ---------- ----- ----------
IOPS Test 1 Seq Read 4K 815K 3.18 GiB/s
IOPS Test 2 Rand Read 4K 100K 391 MiB/s
IOPS Test 3 Seq Write 4K 662K 2.59 GiB/s
IOPS Test 4 Rand Write 4K 79.1K 309 MiB/s
FIO THROUGHPUT TEST COMMANDS:
fio --rw=read --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=1m --iodepth=32 --numjobs=4 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
fio --rw=randread --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=1m --iodepth=32 --numjobs=4 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
fio --rw=write --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=1m --iodepth=32 --numjobs=4 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
fio --rw=randwrite --refill_buffers --norandommap --randrepeat=0 --ioengine=io_uring --bs=1m --iodepth=32 --numjobs=4 --size=128g --time_based --runtime=60 \
--group_reporting --filename=/mnt/vdb/fio-test --name=jobs-test-job --direct=1 --cpus_allowed_policy=split
FIO THROUGHPUT TEST RESULTS:
Test Type Operation Block Size IOPS Bandwidth
----------- ---------- ---------- ----- ----------
Throughput 1 Seq Read 1M 20.3K 19.9 GiB/s
Throughput 2 Rand Read 1M 8039 8.04 GiB/s
Throughput 3 Seq Write 1M 3460 3.46 GiB/s
Throughput 4 Rand Write 1M 1870 1.87 GiB/s
A benchmark result should answer a specific engineering question. Peak throughput, peak IOPS, and lowest latency are different objectives and are often achieved with different block sizes, queue depths, client counts, and access patterns.
When performance stops scaling, determine which resource reached saturation first. A flat storage result does not necessarily indicate a storage-media limit. CPU utilization, PCIe bandwidth, front-end network bandwidth, cluster-network traffic, NIC queue distribution, client capability, metadata service load, or protection/recovery traffic may be the limiting resource.
Increasing queue depth and concurrency can raise aggregate throughput while also increasing latency. For transactional, database, virtual-desktop, and interactive workloads, publish latency together with IOPS or throughput and include percentile latency when available.
Cluster recovery, rebalancing, data reconstruction, multi-site synchronization, and other resilience activities consume CPU, storage, and network resources. Measure these conditions separately when the objective is to understand performance during failure or recovery.
Results should be compared only when the relevant variables are held constant. Changing block size, queue depth, protection policy, client count, network path, caching, encryption, working-set size, or application protocol can materially change the result and invalidate a direct comparison.
The StratiSYSTEM Storage Performance Testing Framework provides a structured method for measuring the complete storage data path rather than relying on isolated peak numbers. Perfprofiler and FIO characterize block and file I/O behavior, IOR evaluates parallel and distributed scaling, Frametest exercises sustained media workloads, S3 Warp measures object-service performance, and iperf3 establishes whether the network is capable of carrying the expected storage traffic.
The most useful results are produced when testing mirrors the intended workload and when the environment is fully documented. CPU and NUMA placement, network topology, storage media, protection policy, client concurrency, protocol configuration, caching, encryption, and cluster state should all be treated as part of the benchmark.
Used consistently, this framework can evaluate StratiSTOR data services, StratiSERV virtual-machine storage paths, HyperSERV HCI, Lustre and HPC environments, S3 object services, high-speed NVMe-oF and iSCSI block storage, SMB/NFS file services, and modern AI or data-intensive infrastructure while producing results that remain understandable and reproducible over time.