AI Kernel Engineer

Active

Job description

Description About Majestic LabsMajestic Labs is a deep-tech pioneer re-architecting AI systems for the data center to systematically dismantle the industry's most critical barrier: the "memory wall." Backed by a $100 million Series A closed in late 2025. We’re a fast-moving AI startup building next-generation infrastructure for the world’s most demanding AI workloads. Our mission is to accelerate the future of intelligence by delivering cutting-edge server solutions optimized for large-scale inference and training. Backed by top-tier investors and led by industry veterans, we’re scaling rapidly and looking for a dynamic partnerships leader to help us win the trust of enterprise and hyperscaler customers.We are a US-Israeli company looking for builders-engineers and operators who have shipped production hardware and know the crucial difference between compelling research and systems that operate at massive scale.If you want to build infrastructure that fundamentally changes what the world can accomplish with AI, come join us!About the positionDesign and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution.Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA).Collaborate with compiler and runtime teams to integrate kernels into Triton and PyTorch pipelines.Profile and tune kernels using tools like Perfetto, VTune, Tracy, Nsight, or custom simulators.Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, MXFP/FP4, etc.).Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team. Requirements Degree at any level in Computer Science, Computer Engineering, or a related field from a recognized university.Strong background in parallel programming (CUDA, Triton, RISC-V Vector (RVV), POSIX Threads, or OpenMP).Deep understanding of memory layout, vectorization, thread/block scheduling, and cache behavior.Experience programming wide-SIMD vector units and matrix/systolic engines.Familiarity with the parallel-computing ecosystem and its high-performance libraries andprimitives, such as BLAS/BLIS, dense linear algebra, and parallel scan/reduce/sort.Skilled in performance analysis and parallel debugging using tools such as GNU Debugger and Nsight.Hands-on experience profiling and optimizing compute or AI workloads (e.g., GEMM, softmax, attention).Solid grasp of numerical stability, precision formats, and mixed precision arithmetic.Proficiency in C++17 or higher, with strong knowledge of standard algorithms, data structures, and generic programming paradigms.Collaborative work style with the ability to operate effectively in multicultural, cross-disciplinary environments.Preferred qualificationsExperience with optimization of irregular algorithms, such as graph computations or sparsenumerical linear algebra, combining high-level data structure design with low-level SIMD and synchronization optimizations.Familiarity with LLVM/MLIR.RISC-V systems or bare-metal programming, including RISC-V vector and matrix extensions.We seek smart, curious, self-driven individuals who enjoy creative problem-solving and continuous learning. Even if you don't meet all requirements, we'd love to hear from you if you have the mindset to tackle complex challenges.Majestic Labs is an equal opportunity employer deeply committed to diversity of background and thought. If you are ready to dare greatly, learn from mistakes, engage in vigorous debates, and help create a step-function shift in human technical capability, we want to meet you!Description About Majestic LabsMajestic Labs is a deep-tech pioneer re-architecting AI systems for the data center to systematically dismantle the industry's most critical barrier: the "memory wall." Backed by a $100 million Series A closed in late 2025. We’re a fast-moving AI startup building next-generation infrastructure for the world’s most demanding AI workloads. Our mission is to accelerate the future of intelligence by delivering cutting-edge server solutions optimized for large-scale inference and training. Backed by top-tier investors and led by industry veterans, we’re scaling rapidly and looking for a dynamic partnerships leader to help us win the trust of enterprise and hyperscaler customers.We are a US-Israeli company looking for builders-engineers and operators who have shipped production hardware and know the crucial difference between compelling research and systems that operate at massive scale.If you want to build infrastructure that fundamentally changes what the world can accomplish with AI, come join us!About the positionDesign and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution.Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA).Collaborate with compiler and runtime teams to integrate kernels into Triton and PyTorch pipelines.Profile and tune kernels using tools like Perfetto, VTune, Tracy, Nsight, or custom simulators.Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, MXFP/FP4, etc.).Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team.Description About Majestic LabsMajestic Labs is a deep-tech pioneer re-architecting AI systems for the data center to systematically dismantle the industry's most critical barrier: the "memory wall." Backed by a $100 million Series A closed in late 2025. We’re a fast-moving AI startup building next-generation infrastructure for the world’s most demanding AI workloads. Our mission is to accelerate the future of intelligence by delivering cutting-edge server solutions optimized for large-scale inference and training. Backed by top-tier investors and led by industry veterans, we’re scaling rapidly and looking for a dynamic partnerships leader to help us win the trust of enterprise and hyperscaler customers.We are a US-Israeli company looking for builders-engineers and operators who have shipped production hardware and know the crucial difference between compelling research and systems that operate at massive scale.If you want to build infrastructure that fundamentally changes what the world can accomplish with AI, come join us!About the positionDesign and implement high-performance compute kernels for AI primitives such as GEMM, attention, normalization, and convolution.Optimize for throughput, latency, and memory hierarchy across heterogeneous compute units (SIMD, matrix engines, DMA).Collaborate with compiler and runtime teams to integrate kernels into Triton and PyTorch pipelines.Profile and tune kernels using tools like Perfetto, VTune, Tracy, Nsight, or custom simulators.Prototype and evaluate precision formats (FP16/BF16/FP8/e5m2, MXFP/FP4, etc.).Contribute to micro-architecture feedback loops, helping co-design ISA and memory features with the hardware team.