{"id":31200,"date":"2026-09-14T00:18:34","date_gmt":"2026-09-13T21:18:34","guid":{"rendered":"https:\/\/hgpu.org\/?p=31200"},"modified":"2026-09-14T00:18:34","modified_gmt":"2026-09-13T21:18:34","slug":"stencil-computation-at-the-intersection-of-ai-and-hpc","status":"publish","type":"post","link":"https:\/\/hgpu.org\/?p=31200","title":{"rendered":"Stencil Computation at the Intersection of AI and HPC"},"content":{"rendered":"<p>Tensor compilers such as TinyTC and OpenAI Triton were originally developed for AI workloads, but the same tiling and memory abstractions can be applied to implement efficient high-order stencils for scientific and industrial applications. We demonstrate this for an 8th-order, 25-point acoustic stencil with boundary conditions over an a demanding-sized grid, targeting GPGPUs, where we compare the hardware-specialized TinyTC implementation with a portable PyTorch\/Triton implementation. The target platforms for evaluation include Intel B70, B580, GPU MAX 1550, NVIDIA A100\/RTX6000 Blackwell\/H100, and AMD MI325x. For instance, on Battlemage B580 TinyTC reaches 15.6 Gpts\/s versus 13.5 Gpts\/s for PT\/Triton under random initialization, while zero-initialized runs reach up to 35.8 Gpts\/s due to hardware memory compression. Using roofline and memory-hierarchy profiling, we show that -as expected- performance is predominantly bandwidth-limited and that compiler-managed L1\/LSC caching can effectively replace programmer-managed shared-memory staging for this stencil class. Overall, the results position TinyTC as the performance-oriented path on Intel hardware and PyTorch\/Triton as a strong portability\/productivity baseline for cross-vendor HPC stencil development.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Tensor compilers such as TinyTC and OpenAI Triton were originally developed for AI workloads, but the same tiling and memory abstractions can be applied to implement efficient high-order stencils for scientific and industrial applications. We demonstrate this for an 8th-order, 25-point acoustic stencil with boundary conditions over an a demanding-sized grid, targeting GPGPUs, where we [&hellip;]<\/p>\n","protected":false},"author":351,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[11,3],"tags":[1733,1438,2186,1782,1682,2204,2207,2133,20,2066,2132,2200,1586,1728,1845,2182],"class_list":["post-31200","post","type-post","status-publish","format-standard","hentry","category-computer-science","category-paper","tag-ai","tag-amd","tag-amd-radeon-instinct-mi325x","tag-computer-science","tag-hpc","tag-intel-arc-b580","tag-intel-b70","tag-intel-data-center-gpu-max-1550","tag-nvidia","tag-nvidia-a100","tag-nvidia-h100","tag-nvidia-rtx-pro-6000","tag-performance-portability","tag-stencil-computation","tag-sycl","tag-triton"],"views":375,"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/hgpu.org\/index.php?rest_route=\/wp\/v2\/posts\/31200","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/hgpu.org\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/hgpu.org\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/hgpu.org\/index.php?rest_route=\/wp\/v2\/users\/351"}],"replies":[{"embeddable":true,"href":"https:\/\/hgpu.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=31200"}],"version-history":[{"count":0,"href":"https:\/\/hgpu.org\/index.php?rest_route=\/wp\/v2\/posts\/31200\/revisions"}],"wp:attachment":[{"href":"https:\/\/hgpu.org\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=31200"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/hgpu.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=31200"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/hgpu.org\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=31200"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}