CRPL Earns Best Poster Award at IEEE Cluster 2026
2026


The Computational Research and Programming Languages (CRPL) group had a strong showing at IEEE Cluster 2026, with three research posters accepted for presentation. Among the three accepted works, the poster presented by Zack Sollenberger, Rahul Patel, and Saieda Ali Zada received the Best Poster Award. Together, the projects showcased CRPL research spanning compiler validation, GPU memory systems, scientific computing, and AI-assisted optimization for high-performance computing.

Zack Sollenberger, Rahul Patel, and Saieda Ali Zada presented the group’s Best Poster Award-winning work, which focuses on creating a dataset of valid and invalid OpenACC and OpenMP compiler tests. The project addresses an important challenge in compiler validation and verification: while valid examples of directive-based programs are widely available, intentionally invalid compiler tests designed to exercise specific errors and specification violations can be much more difficult to find.
To build the dataset, the team utilized a Retrieval-Augmented Generation (RAG) workflow to extract relevant information from OpenACC and OpenMP documentation and associate specification information with individual compiler tests. The team also developed 20 unique issue types that could be systematically injected into valid compiler tests to generate corresponding invalid cases. This work provides a structured dataset that can support compiler testing, specification compliance, diagnostic evaluation, and future research into AI-assisted compiler validation. The quality and impact of this work were recognized at IEEE Cluster 2026 with the Best Poster Award.
Nathan Graddon presented research on characterizing CUDA Unified Memory on NVIDIA Grace–Blackwell GPUs using STREAM microbenchmarks. The work investigates how the location and movement of data within a unified memory system can significantly affect application performance, even when the GPU computation itself remains unchanged.
Using a STREAM Triad kernel, the study evaluated five different memory placement and management scenarios while keeping the computational workload consistent. Application-level useful memory bandwidth was then measured across each case to isolate the performance effects associated with where data resides and how it reaches the GPU. The results highlight an important distinction in heterogeneous computing: although CUDA Unified Memory can simplify programming by providing CPUs and GPUs with access to shared allocations, unified addressability does not guarantee uniform performance. Data placement, migration, reuse, and the underlying memory path continue to play an important role in determining achievable GPU performance.
Connor Vitz presented research exploring the use of OpenEvolve, an LLM-guided evolutionary search framework, to tune Ozaki Scheme GPU kernels. The Ozaki Scheme enables higher-accuracy computation using lower-precision INT8 and FP8 tensor cores, offering the potential to emulate FP64-level numerical accuracy while taking advantage of the much higher throughput available from modern GPU tensor cores.
The project investigated whether an AI-guided optimization system could automatically identify effective kernel configurations and reduce the amount of manual performance tuning required. The experiments demonstrated real throughput advantages for the emulated approach compared with native FP64 computation, while OpenEvolve successfully converged on an acceptable configuration. Evaluation on a real scientific workload also revealed areas where additional optimization is needed, providing a direction for continued development and future research.
Together, the three posters demonstrate the breadth of research being conducted within CRPL—from compiler validation and directive-based programming to emerging GPU memory architectures and AI-assisted scientific computing. The group’s three accepted posters and Best Poster Award at IEEE Cluster 2026 represent an exciting recognition of CRPL’s continued work in high-performance computing, programming models, compilers, and GPU systems.
