Hewlett Packard Enterprise
High Performance Compute Systems Site Lead (Onsite - LANL)
All, New Mexico, United States of America
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 27 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
High Performance Compute Systems Site Lead (Onsite - LANL)
This role has been designated as ‘Remote/Teleworker’, which means you will primarily work from home.
Who We Are:
Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.
Job Description:
Join Hewlett Packard Enterprise as a Technical Site Lead supporting some of the world’s most advanced high-performance computing (HPC) and AI systems at Los Alamos National Laboratory (LANL), including large-scale Linux and HPE Cray environments. This senior, hands-on individual contributor position provides onsite technical and service-delivery leadership for HPE hardware engineers, Linux system administrators, software analysts, inventory specialists, and remote engineering resources supporting mission-critical HPE Cray and related platforms that enable scientific discovery and national security.
This is a technical leadership role with no direct reports. The Technical Site Lead combines hands-on Linux systems administration and HPC support experience with enterprise server hardware capability, structured troubleshooting, customer-facing leadership, incident coordination, and disciplined service delivery. The role is accountable for coordinating the onsite team's technical execution, maintaining operational readiness and service quality, managing escalations, planning maintenance, supporting major incidents, and preparing the site for next-generation HPC and AI platforms.
Operating as the site's technical leader, this individual establishes daily priorities, coordinates assignments, provides clear technical direction, mentors team members, follows through on commitments, and removes barriers affecting service delivery. Success requires the ability to lead through influence, work effectively across onsite and remote engineering organizations, communicate clearly with customer and HPE stakeholders, and maintain consistent execution in a complex, mission-critical computing environment.
US Citizenship and the ability to obtain and maintain a DOE Q Clearance are required.
Daily onsite work is required in Los Alamos, New Mexico. This is not a remote or hybrid position. Standard work hours are Monday through Friday, either 8:00 a.m. to 5:00 p.m. or 7:00 a.m. to 4:00 p.m., with additional onsite support required for planned maintenance, major incidents, and on-call lead responsibilities.
Key Responsibilities
Technical Leadership & Service Delivery
- Provide day-to-day technical leadership and technical guidance to onsite HPE hardware engineers, Linux system administrators, and software analysts, while coordinating work across inventory specialists and remote engineering resources supporting large-scale HPE Cray EX and related HPC and AI systems.
- Set and communicate daily and weekly technical priorities based on system health, open support cases, scheduled maintenance, customer priorities, operational risk, service commitments, and available staffing.
- Serve as the primary onsite technical focal point for day-to-day system support, technical escalations, maintenance activities, and service-delivery risks, partnering with the DSM on customer governance, executive escalations, and broader service-delivery matters.
- Maintain a current view of system health, support-case status, technical risks, maintenance actions, ownership, and outstanding commitments. Prepare and lead routine onsite operational reviews with the customer and onsite team in coordination with the DSM and HPC leadership.
- Ensure support cases contain accurate technical details, diagnostic evidence, business impact, troubleshooting history, current ownership, and clearly defined next actions.
- Drive timely escalation through established HPE processes and help team members engage next-tier support, engineering, product teams, and other resources needed to advance diagnosis and resolution.
- Plan and coordinate system upgrades, maintenance windows, installations, acceptance activities, and other production-impacting work in accordance with customer change-control requirements and HPE support processes. Review prerequisites, risks, execution steps, and expected outcomes with the onsite team and customer before work begins.
- Confirm required staffing, parts, tools, test equipment, documentation, communications, escalation contacts, rollback plans, and contingency coverage are in place for planned work.
- Lead the HPE onsite response during major incidents by aligning technical priorities with the customer-designated incident lead and DSM, organizing HPE resources, maintaining clear communications, tracking actions and decisions, and ensuring appropriate escalation.
- Coordinate HPE root-cause analysis, corrective actions, lessons learned, and documentation updates after significant incidents or recurring issues involving HPE-supported hardware, firmware, and infrastructure.
- Identify risks to system availability, service-level commitments, maintenance schedules, operational readiness, or customer satisfaction and escalate them promptly to the DSM.
- Coordinate onsite parts inventory, repair materials, tools, test equipment, and other HPE-owned resources using approved HPE business systems and controls.
Operational & Team Support
- Build and maintain effective, professional working relationships with onsite team members, remote engineering organizations, HPE leadership, customer technical staff, and customer management.
- Facilitate regular team coordination discussions to organize work, confirm ownership, review progress, surface technical blockers, and ensure commitments are completed or appropriately escalated.
- Provide technical mentoring, coaching, and practical guidance to team members without assuming formal people-management authority.
- Promote a culture of accountability, disciplined troubleshooting, accurate documentation, safe work practices, knowledge sharing, and professional customer engagement.
- Track completion of required HPE, customer, security, safety, and technical training; notify team members of approaching deadlines and escalate overdue or at-risk requirements to the DSM.
- Coordinate onsite coverage using approved team schedules and planned leave information. Identify and escalate potential gaps in business-hours, maintenance, or on-call coverage before they affect service delivery.
- Provide the DSM with fact-based observations regarding technical performance, development needs, recognition opportunities, and issues requiring formal management attention.
- Coordinate site-specific onboarding and operational readiness for new team members, including required accounts, badges, site access, training, workspace, equipment, and introductions to key stakeholders.
- Maintain accurate site procedures, contact lists, escalation paths, team schedules, operational references, and other information required for effective day-to-day support.
- Maintain a clean, safe, secure, and organized working environment in accordance with HPE and customer requirements.
- Participate in the on-call rotation and provide additional onsite support when required for 24x7 operations, planned maintenance, system outages, and major incidents.
Hands-on Technical Contribution
Use Linux command-line and diagnostic tools in Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or comparable environments
Similar jobs
- F-35 Training Systems Site Activation Analyst, SeniorBah · 6 LocationsFirst seen 8d ago
- High Performance Compute Systems Site Lead (Onsite - LANL)Hpe · All, New Mexico, United States of AmericaFirst seen 5d ago
- High Performance Compute Systems Site Lead (Onsite - LANL)Hpe · All, New Mexico, United States of AmericaFirst seen 5d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job