
Staff Test Engineer - AI SW Stack ( 9-13 years)
Sandisk understands how people and businesses consume data and we relentlessly innovate to deliver solutions that enable today’s needs and tomorrow’s next big ideas. With a rich history of groundbreaking innovations in Flash and advanced memory technologies, our solutions have become the beating heart of the digital world we’re living in and that we have the power to shape.
Sandisk meets people and businesses at the intersection of their aspirations and the moment, enabling them to keep moving and pushing possibility forward. We do this through the balance of our powerhouse manufacturing capabilities and our industry-leading portfolio of products that are recognized globally for innovation, performance and quality.
Sandisk has two facilities recognized by the World Economic Forum as part of the Global Lighthouse Network for advanced 4IR innovations. These facilities were also recognized as Sustainability Lighthouses for breakthroughs in efficient operations. With our global reach, we ensure the global supply chain has access to the Flash memory it needs to keep our world moving forward.
We are seeking an experienced AI Software Stack Test Lead to define, drive, and execute validation strategy for a next-generation AI computational storage platform. This role owns end-to-end verification of the AI software stack, including neural network operators, compiler-generated workloads, runtime execution, memory management, scheduling systems, and model-level performance.
The ideal candidate combines strong software validation expertise with deep knowledge of AI/ML execution pipelines and will lead validation efforts across compiler, runtime, kernel, and hardware teams to ensure correctness, scalability, reliability, and performance.
Key Responsibilities
Validation Strategy & Technical Leadership
- Own and define the end-to-end validation strategy for the AI software stack.
- Establish validation methodologies, coverage metrics, quality gates, and release-readiness criteria.
- Lead validation efforts across operator libraries, runtime systems, execution frameworks, and AI workloads.
- Mentor engineers and drive best practices for automation, performance validation, and debugging.
Functional Validation
- Validate neural network operators and compute kernels
- Validate graph-level execution and end-to-end model inference workflows.
- Verify numerical correctness against reference frameworks such as PyTorch or TensorFlow.
- Ensure behavioral consistency across software releases and hardware revisions.
Runtime & Execution Validation
- Validate memory allocation and movement across host and device environments.
- Execute stress, concurrency, and multi-device validation scenarios.
Performance & Scalability Validation
- Design and maintain benchmarking frameworks for:
- Operator-level performance
- Model-level performance
- Latency and throughput
- Time-to-first-token/result
- Multi-device scaling efficiency
Required Qualifications
- 10+ years of experience in AI software validation, systems software validation, performance engineering, or related areas.
- Strong C++ software development expertise.
- Strong Object-Oriented Programming fundamentals.
- Advanced Python programming for automation and test infrastructure.
- Deep understanding of:
- AI/ML execution pipelines
- Neural network operators
- Deep learning frameworks (PyTorch preferred)
- Runtime architectures
- Scheduling systems
- Parallel execution models
- Memory hierarchy and DMA concepts
- Experience validating large-scale AI workloads and inference systems.
- Proven experience building automated test and benchmarking frameworks.
- Strong debugging and root-cause analysis skills.
Preferred Qualifications
- Experience with MLIR, XLA, StableHLO, TVM, ONNX Runtime, TensorRT, or similar compiler/runtime stacks.
- Experience validating GPU, NPU, FPGA, or custom accelerator platforms.
- Hands-on experience with profiling tools such as: VTune, Nsight Systems, Nsight Compute, perf, gprof
- Experience with distributed inference and multi-device execution.
- Familiarity with LLM, multimodal, and generative AI workloads.
- Experience leading software quality and validation teams.
Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.
Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at jobs.accommodations@sandisk.com to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.
Similar jobs

Sr. Staff, Software Machine Learning Test Engineer

Sr. SW Test Engineer - Neural Library (5-8 years)

RF Product and Test Engineer, Staff

Sr. Staff Systems Test Engineer

Product Test Engineer - Digital Processors and Sensors characterization, up to Staff
