← Zurück zur Liste
Stelle

Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program

Sonstige • Vor Ort • Vollzeit Israel Kfar Saba, Israel

Sandisk understands how people and businesses consume data and we relentlessly innovate to deliver solutions that enable today’s needs and tomorrow’s next big ideas. With a rich history of groundbreaking innovations in Flash and advanced memory technologies, our solutions have become the beating heart of the digital world we’re living in and that we have the power to shape.

Sandisk meets people and businesses at the intersection of their aspirations and the moment, enabling them to keep moving and pushing possibility forward. We do this through the balance of our powerhouse manufacturing capabilities and our industry-leading portfolio of products that are recognized globally for innovation, performance and quality.

Sandisk has two facilities recognized by the World Economic Forum as part of the Global Lighthouse Network for advanced 4IR innovations. These facilities were also recognized as Sustainability Lighthouses for breakthroughs in efficient operations. With our global reach, we ensure the global supply chain has access to the Flash memory it needs to keep our world moving forward.

Hive is a swarm of hundreds of identical storage nodes (Marvell/XSight DPU + SSD), each running a full Ceph plane: OSD, Monitor (MON), Manager (MGR), MDS, plus SeaStore and the NFS front end. Standing up, re-configuring, and recovering a cluster of this size by hand does not scale. We need an engineer who owns the automated hardware discovery and Ceph role-assignment pipeline: the flow that inventories every node and its hardware, decides which daemons each node should run, and drives the cluster from bare metal to a serving state. The flow must run fully autonomously by default, with a clean manual-override path for lab, bring-up, and failure-injection scenarios.

This is a specialized automation and lab-hardware-discovery discipline that the Hive team does not currently have dedicated ownership for. It is directly on the critical path for every test cluster, every silicon bring-up, and every customer-shaped deployment.

What you'll own

  • Automated hardware discovery. Detect nodes as they power on and enumerate their hardware (DPU/SoC model, SSD media, DRAM, RNIC/network ports, BMC) using out-of-band and in-band inventory (BMC/Redfish/IPMI, PXE/DHCP boot, cloud-init/first-boot agents). Produce a single authoritative machine inventory that the rest of the pipeline consumes.
  • Role assignment and cluster composition. Given the discovered inventory, decide and apply which Ceph roles each node runs — OSD, MON, MGR, MDS (and the NFS gateway) — encoded as declarative placement specifications. Own MON quorum sizing and placement, MGR redundancy, MDS/gateway placement, and CRUSH map / failure-domain layout so data and metadata land correctly across the swarm.
  • Bare-metal → serving bring-up. Drive the end-to-end sequence: node provisioning, OS/image deploy, cluster bootstrap, daemon deployment via the Ceph orchestrator (cephadm-style), ceph-volume-style OSD provisioning on the SSD media, and health convergence to HEALTH_OK.
  • Autonomous and manual modes. Make the default path zero-touch (a rack powers on and self-assembles into a healthy cluster), while exposing deterministic manual controls to pin roles, hold a node out, force a specific topology, or reproduce a customer/lab configuration for testing.
  • Lifecycle and recovery automation. Node add/remove, drain and rebalance, daemon replacement, MON re-quorum after loss, MDS/OSD failover validation, and re-discovery after re-imaging — integrated with Hive's Fast Recovery + BMC work.
  • Reconciliation and drift control. Continuously compare declared desired state against observed cluster state and converge — the same idea Ceph's orchestrator applies to service specs — with clear reporting when reality diverges from intent.
  • CI/lab integration. Wire the pipeline into automated test so any commit can spin up a correctly-composed multi-node cluster on real hardware and tear it down cleanly.

Core (must-have):

  • Deep experience in lab hardware discovery, inventory, and automated provisioning at fleet scale.
  • Bare-metal automation: PXE/DHCP boot, BMC out-of-band management (Redfish/IPMI), image/OS deployment, cloud-init/first-boot.
  • Infrastructure-as-code and config automation (e.g., Ansible/Terraform-class tooling) with a declarative, reconcile-to-desired-state mindset.
  • Strong scripting/automation (Python and shell) and CI systems (Jenkins/GitLab CI or equivalent).
  • Comfort designing systems that are autonomous by default with reliable manual override.

Strongly preferred (or ramp-up expected):

  • Ceph operational knowledge: OSD/MON/MGR/MDS roles, the orchestrator/cephadm model, placement specs, ceph-volume, CRUSH maps and failure domains, MON quorum, and cluster health/lifecycle.
  • Distributed-systems fluency: quorum/Paxos intuition, rebalance/recovery behavior, failure-domain reasoning.
  • Networking bring-up for storage fabrics (RoCEv2/Ethernet, port discovery).
  • Familiarity with DPU/SoC-based nodes and constrained-node environments.

Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at jobs.accommodations@sandisk.com to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Ähnliche Stellen