DDN
Staff Engineer, Lustre
Santa Clara Office
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Lead / management
- Stated salary
- $200,000 – $250,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 2 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
We are seeking a Staff Engineer with 10+ years of experience in distributed storage and Linux-based systems engineering. This is a hands-on senior technical role focused on design, debugging, performance, and operational excellence across LustreFS and adjacent stack components. The ideal candidate brings strong expertise in one or more Lustre subsystems, can independently drive complex investigations, and collaborates effectively across engineering, QE, support and release teams. Engineers who are comfortable using AI to accelerate triage, debugging, code comprehension and new feature design will be especially valuable.
Key Responsibilities
Design, develop and debug LustreFS features, fixes and enhancements across relevant subsystems such as llite, MDS/MDT, OSS/OST, LDLM and LNet.
Investigate customer and scale-related defects, drive root-cause analysis and implement high-quality fixes with strong attention to correctness and maintainability.
Contribute to performance tuning, failure analysis and reliability improvements for large-scale Lustre deployments.
Participate actively in code reviews, design reviews and subsystem discussions, bringing rigor to testing and operational readiness.
Work closely with QE and support to reproduce issues, improve diagnostic data quality and increase coverage for high-risk failure scenarios.
Help document subsystem behavior, debugging approaches, known failure patterns and operational best practices.
Use AI-assisted tools where appropriate to speed up issue triage, summarize logs, improve code understanding and capture reusable lessons learned.
Required Qualifications
10+ years of experience in systems software, distributed systems, storage, Linux kernel or filesystem engineering.
Strong experience in LustreFS development, support or performance engineering with depth in at least one major subsystem.
Strong C programming and Linux systems debugging skills.
Working knowledge of Linux kernel internals, filesystem semantics, networking and performance analysis.
Experience with LNet and/or high-performance transports such as RDMA, InfiniBand, RoCE or TCP-based storage networking.
Ability to debug and resolve issues spanning multiple layers including client, server, network and backend storage.
Strong collaboration skills and the ability to work across functions in a fast-moving engineering environment.
Preferred Skills
Experience in HPC, AI infrastructure or large-scale parallel storage environments.
Exposure to metadata-heavy and throughput-heavy workload characterization and tuning.
Familiarity with ZFS, ldiskfs, NVMe-backed storage and related observability / performance tooling.
Experience creating test plans, reproducer frameworks, runbooks or diagnostic automation.
Comfort using AI tools to accelerate debugging, code reviews, triage, documentation and early-stage design ideation.
Experience mentoring junior engineers or leading focused technical efforts within a subsystem.
What You Will Work On
Hands-on development and debugging of LustreFS defects, performance issues and subsystem enhancements.
Customer-facing and scale-related issue investigation across llite, metadata, object storage, LNet and transport layers.
Collaborative design and implementation of reliability, observability and serviceability improvements.
Reviewing and validating fixes through targeted tests, failure injection, log analysis and performance characterization.
Using AI-assisted workflows to accelerate triage, debug loops, code understanding and documentation quality.
Contributing to team redundancy by strengthening documentation, code review quality and subsystem knowledge sharing.
Why This Role Matters
This role is central to building durable engineering redundancy in LustreFS: expanding deep subsystem ownership, reducing concentration risk, and accelerating next-generation delivery through strong engineering fundamentals and AI-enabled execution.
Similar jobs
- Information Security Engineer, PrincipalBSC · El Dorado Hills, CA, United States; WA, United States; RI, United States; OH, United States; MO, United States; CA, United States; Long Beach, CA, United States; AZ, United States; CO, United States; FL, United States; GA, United States; MD, United States; MN, United States; NV, United States; OR, United States; Lodi, CA, United States; Rancho Cordova, CA, United States; San Diego, CA, United States; AL, United States; IL, United States; VA, United States; WI, United States; TX, United States; NY, United StatesFirst seen today
- Information Security Engineer, PrincipalBSC · El Dorado Hills, CA, United States; CA, United States; Long Beach, CA, United States; Lodi, CA, United States; Oakland, CA, United States; Rancho Cordova, CA, United States; San Diego, CA, United StatesFirst seen today
- Lead QA Engineer - Tax Product DevelopmentBDO USA Experienced · New York, NY, United States; Atlanta, GA, United States; Austin, TX, United States; Baltimore, MD, United States; Boston, MA, United States; Charlotte, NC, United States; Cherry Hill, NJ, United States; Chicago, IL, United States; Cincinnati, OH, United States; Cleveland, OH, United States; Columbus, OH, United States; Dallas, TX, United States; Detroit, MI, United States; Fort Lauderdale, FL, United States; Fort Worth, TX, United States; Grand Rapids, MI, United States; Greenville, SC, United States; Houston, TX, United States; Indianapolis, IN, United States; Jacksonville, FL, United States; Kalamazoo, MI, United States; Melville, NY, United States; Madison, WI, United States; McLean, VA, United States; Memphis, TN, United States; Miami, FL, United States; Milwaukee, WI, United States; Minneapolis, MN, United States; Nashville, TN, United States; Norfolk, VA, United States; Oak Brook, IL, United States; Omaha, NE, United States; Orlando, FL, United States; Philadelphia, PA, United States; Pittsburgh, PA, United States; Potomac, MD, United States; Raleigh, NC, United States; Rosemont, IL, United States; Richmond, VA, United States; St Louis, MO, United States; Stamford, CT, United States; Tampa, FL, United States; Tulsa, OK, United States; Washington, DC, United States; West Palm Beach, FL, United States; Wilmington, DE, United States; Woodbridge, NJ, United StatesFirst seen today
- Senior Staff Embedded Software EngineerFord Global · Long Beach, CA, United StatesFirst seen today
- Principal Software Engineer, Core InfrastructureOracle · Nashville, TN, United StatesFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job