GPGPU-SIM – Cycle-Level Simulator for Nvidia GPUs with CUDA or OpenCL Workloads
Discussion | hackernews | Author: peter_d_sherman

GPGPU-SIM – Cycle-Level Simulator for Nvidia GPUs with CUDA or OpenCL Workloads
Discussion | hackernews | Author: peter_d_sherman
Show HN: Rgpu – a PyTorch device whose tensors live on a remote GPU
Discussion | hackernews | Author: boxstream
Unroll a loop, make shader 10x faster
Discussion | lobsters | Author: calvin
Efficient FlashAttention on Blackwell via Fixed-Shift Softmax and Persistent Scheduling
tech_blogs_arxiv | Author: Oleksandr Stashuk, Hongtao Yu, Jay Shah
Show HN: Enki – Write GPU compute kernels in pure stable Rust
Discussion | hackernews | Author: enki_runtime
UMAP embeddings can vary between runs because stochastic negative sampling makes the repulsive forces depend on sampling order. ibUMAP addresses this with a coherent, field-based formulation evaluated…
#UMAP #DimensionalityReduction #MachineLearning #GPUComputing
https://arxiv.org/abs/2610.01445
KernelBench: Can LLMs Write GPU Kernels? – Benchmark and Toolkit, Torch –> CUDA
Discussion | hackernews | Author: peter_d_sherman
SoRoMoX is a Python/JAX framework for reduced-order soft-robot simulation, built on Cosserat-rod theory and accelerated with Warp GPU kernels for parallel continuum-model execution. It supports differentiable,…
#SoftRobotics #JAX #Robotics #GPUComputing
https://arxiv.org/abs/2608.06650
New arXiv paper proposes map-conditioned autoregressive generation of human mobility, using a road raster to condition a decoder emitting 31.25 m mesh-cell tokens, tested on trajectories from Ishikawa Prefecture.
#Semiconductors #AIInfrastructure #GPUComputing
https://arxiv.org/abs/2609.32360
🚀 Build practical skills in HPC performance engineering at @it4innovations.bsky.social in Czechia.
Profile, benchmark and optimise parallel applications across modern CPU and GPU systems using MPI, OpenMP, CUDA, Slurm and more.
🔗 Apply Now: hpctrain.eu/traineeship/...
🚀 Interested in AI and HPC?
Join @hpctrain.bsky.social at Research Institute in Sweden and gain hands-on experience optimising and benchmarking large-scale AI applications across CPU and GPU-based HPC systems.
Show HN: Agentic CUDA Kernel Optimizer
Discussion | hackernews | Author: bertaye
fatbin tools for object files
tech_blogs_redplait_blogspot_com | Author: redp
CUDA.jl 6.4 is here 🚀
🟢 Improved NVIDIA Jetson support
🟢 Compiled GPU code caching for faster TTFX
🟢 CUDA 13.4 support
🟢 LLVM 23 upgrade
See what's new → juliagpu.org/post/2026-09...
Two GPU tools for data too big to look at
cuGraph does graph analytics on the GPU; Omniverse does real-time 3D simulation.
#NCAAIIO #AI #DataScience #GPUComputing
Full NCA-AIIO explanation, free: https://navyduck.com/nvidia/ai-infrastructure/nca-aiio/q117-a-research-team-needs-to
CUDA Toolkit 13.4技術解説:Windows on Arm対応とGPU共有リソース管理の進化
CUDA 13.4が登場。Windows on Arm対応と共有GPUの制御強化を詳説。