hirly

GDIT

Senior Data Engineer - Secret Clearance Required

USA VA Sterling

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at GDIT first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included
Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Data & ML
Seniority
Senior
Country
US
Work mode
On-site / unstated
First seen by hirly
9 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Type of Requisition:

Regular

Clearance Level Must Currently Possess:

Secret

Clearance Level Must Be Able to Obtain:

Secret

Public Trust/Other Required:

None

Job Family:

Data Science and Data Engineering

Job Qualifications:

Skills:

Artificial Intelligence (AI), Data Solutions, Technical Solutions Certifications:

None Experience:

10 + years of related experience US Citizenship Required:

Yes

Job Description:

Own your opportunity to turn data into measurable outcomes for our customers’ most complex challenges. As a Data Engineer Sr Principal at GDIT, you’ll power innovation to drive mission impact and grow your expertise to power your career forward.

The ideal candidate will have strong experience with relational databases, data pipelines, ETL/ELT processes, Python development, and enterprise data integration. The candidate should also be comfortable working with containerized applications, Linux environments, and emerging AI technologies, including vector databases and Retrieval-Augmented Generation (RAG).

This individual will collaborate with software engineers, AI/ML engineers, cybersecurity professionals, database administrators, and enterprise architects to deliver scalable, secure, and reliable data solutions within highly controlled government environments.

SECRET CLEARANCE REQUIRED

Key Responsibilities

1. Chatbot – AI Platform Data Engineering

  • Design, develop, and maintain the data infrastructure supporting our chatbot and related AI-powered applications.
  • Manage and optimize PostgreSQL databases, including pgvector extensions for vector similarity search.
  • Develop and maintain data ingestion pipelines for structured, semi-structured, and unstructured data.
  • Integrate document processing technologies such as Docling and Tesseract into automated data pipelines.
  • Support document metadata extraction, transformation, validation, indexing, and storage.
  • Collaborate with AI/ML engineers to optimize data retrieval, embedding storage, and RAG workflows.
  • Develop Python-based services, scripts, and APIs to facilitate data movement between systems.
  • Support the migration of database services and supporting components into Podman containers.
  • Improve database reliability, performance, observability, backup procedures, and recoverability.
  • Troubleshoot data integration issues across application services, databases, and AI inference systems.

2. Enterprise Data Integration & Data Fabric

  • Assist in evaluating and designing a potential enterprise data fabric architecture.
  • Develop solutions for integrating data across disparate applications, databases, and enterprise systems.
  • Design reusable data pipelines and integration patterns that support multiple applications and business domains.
  • Support the development of enterprise data catalogs, metadata management, data lineage, and data discovery capabilities.
  • Evaluate approaches for data virtualization, data federation, and distributed data access.
  • Define and implement data quality, consistency, validation, and governance controls.
  • Collaborate with enterprise architects and stakeholders to identify integration requirements and technical solutions.
  • Help establish common data models, integration standards, and reusable data services.
  • Evaluate open-source and commercially available data integration technologies for use in secure government environments.

3. Data Migration & Modernization

  • Plan and execute data migrations between legacy systems, modern databases, and enterprise applications.
  • Assess source systems, database schemas, data quality, dependencies, and migration requirements.
  • Develop ETL/ELT pipelines using Python, SQL, and appropriate data integration frameworks.
  • Perform data extraction, transformation, cleansing, mapping, reconciliation, and validation.
  • Design repeatable migration processes that minimize operational risk and data loss.
  • Support migration testing, performance optimization, rollback planning, and post-migration verification.
  • Develop scripts and automation to streamline recurring data migration activities.
  • Produce technical documentation, migration plans, data mapping specifications, and operational procedures.

4. Infrastructure, Automation & Security

  • Develop and support containerized data services using Podman and container orchestration tools.
  • Support applications and databases running on Red Hat Enterprise Linux and Windows Server.
  • Automate database deployment, configuration, monitoring, and maintenance activities.
  • Implement secure data access controls, encryption, audit logging, and data handling practices.
  • Support deployment of software and data services in air-gapped and classified environments.
  • Collaborate with cybersecurity and infrastructure teams to ensure compliance with applicable federal security requirements.
  • Participate in vulnerability remediation, software evaluation, change management, and accreditation support activities.
  • Implement monitoring and operational logging solutions using tools such as Splunk.

Required Qualifications

  • Bachelor's degree in Computer Science, Data Engineering, Information Systems, Software Engineering, or a related technical field, or equivalent practical experience.
  • 7+ years of experience in data engineering, database development, data integration, or related software engineering disciplines.
  • Strong proficiency in Python for data processing, scripting, automation, and service integration.
  • Advanced SQL skills, including query optimization, database schema design, and performance tuning.
  • Experience administering, developing against, or optimizing PostgreSQL or comparable enterprise relational database systems.
  • Demonstrated experience designing and implementing ETL/ELT pipelines.
  • Experience with enterprise data migrations, including data mapping, transformation, validation, and reconciliation.
  • Familiarity with REST APIs, JSON, structured and semi-structured data formats.
  • Experience with Linux administration, shell scripting, and deployment automation.
  • Understanding of data modeling, relational databases, data integrity, and data lifecycle management.
  • Experience with Git, automated testing, and modern software development practices.
  • Ability to design, document, and troubleshoot complex data workflows.
  • Strong communication skills and ability to collaborate across multidisciplinary engineering teams.

Preferred Qualifications

  • Experience supporting AI/ML platforms, LLM applications, or Retrieval-Augmented Generation solutions.
  • Familiarity with vector databases, PostgreSQL pgvector, embeddings, and semantic search.
  • Experience with document ingestion and content extraction technologies such as Docling, Apache Tika, or Tesseract.
  • Experience with container technologies such as Podman or Docker.
  • Familiarity with Redis and PgBouncer.
  • Experience with enterprise data fabric, data mesh, or data virtualization architectures.
  • Experience with Apache Airflow, Apache NiFi, or similar data pipeline orchestration frameworks.
  • Experience with Apache Kafka or other event-driven data integration technologies.
  • Familiarity with data catalogs, metadata management, data lineage, and data governance frameworks.
  • Experience with cloud-based data platforms or hybrid data architectures, including AWS or Azure.
  • Experience modernizing legacy databases and enterprise data platforms.
  • Familiarity with secure, disconnected, air-gapped, or classified IT environments.
  • Knowledge of federal cybersecurity standards, RMF processes, and Authority to Operate (ATO) requirements.
  • Experience with Splunk, automated testing, and CI/CD practices.

The likely salary range for this position is $153,000 - $207,000. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.

Scheduled Weekly Hours:

40

Travel Required:

Original posting on GDIT's site ↗

Listed on hirly, a job board. hirly is not the employer: GDIT is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job