FlexGroup and FlexCache Optimization for AWS-Based CAE/HPC File Workloads: A Latency and Throughput Evaluation
Keywords:
computer-aided engineering, Amazon FSx for NetApp ONTAP, FlexCache, FlexGroupAbstract
Background: Computer-aided engineering (CAE) and high-performance computing (HPC) workflows place conflicting demands on shared file storage. Large sequential solver I/O, metadata-intensive preprocessing, repeated reads of common models, bursty checkpointing, and geographically separated data sources can expose both latency and aggregate-throughput limits. NetApp FlexGroup and FlexCache provide complementary mechanisms for addressing these limits in Amazon Web Services (AWS)-based ONTAP environments.
Objective: To evaluate the performance mechanisms, workload dependencies, and optimization strategies by which FlexGroup and FlexCache can improve latency and throughput for AWS-based CAE/HPC file workloads.
Methods: A structured narrative review synthesized peer-reviewed systems literature, HPC I/O characterization studies, NetApp technical reports, and AWS technical evaluations published through 2023. Evidence was organized by storage abstraction, workload pattern, metric, and optimization mechanism. Quantitative findings were retained only when directly reported; otherwise, engineering relationships were expressed analytically.
Results: FlexGroup improves scale by distributing files and metadata across constituent volumes while maintaining a single namespace. In a published evaluation, FlexGroup achieved normalized throughput of 0.90 for the SPEC SFS EDA profile relative to a local FlexVol baseline of 1.00, compared with 0.78 for a remote FlexVol placement; scaling across eight nodes reached 5.28 times the one-node EDA throughput. FlexCache addresses a different bottleneck: distance to data. Repeated reads are served from a sparse local cache, whereas first-read misses and write-around operations remain dependent on network and origin response time. AWS guidance therefore emphasizes cache-hit ratio, working-set sizing, prewarming, and grouping clients that process similar datasets.
Conclusions: FlexGroup and FlexCache should not be treated as interchangeable acceleration features. FlexGroup is primarily a scale-out and load-distribution mechanism; FlexCache is primarily a data-locality mechanism. The strongest CAE/HPC design combines workload-aware FlexGroup sizing with FlexCache placement close to compute, explicit cold-versus-warm benchmarking, and continuous observation of cache-hit ratio, network latency, origin response time, constituent balance, throughput, and IOPS.
Downloads
Published
Issue
Section
License
Copyright (c) 2023 Mohammed Nazir (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Authors retain copyright in their published work.
Articles published by the International Journal of Business & Computational Sciences (IJBCS) are distributed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License (CC BY-NC-ND 4.0).
Under this license, users may copy and redistribute the published material in any medium or format for non-commercial purposes, provided that appropriate credit is given to the author(s) and the International Journal of Business & Computational Sciences, a link to the license is provided, and the material is not modified, adapted, remixed, transformed, or built upon.
The full terms of the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License are available at: