Data Engineer, Databricks (Mid-Level)
- Job Title
- Data Engineer, Databricks (Mid-Level)
- Job ID
- 27785433
- Work From Home
- Yes
- Work Remote
- Yes
- Location
- US Work From Home Remote
- Other Location
- Description
-
Full Tilt Data is a trusted data, analytics, and IT consulting firm specializing is health related services for the federal government. Established in 2023 by a group of founders that bring 15+ years of industry experience. We are passionate about harnessing the power of data through our comprehensive data management solutions. We mobilize the right people, skills, and technologies to help all types of organizations and companies improve their performance and data management.
Position Summary
We are looking for a Data Engineer, Databricks (Mid-Level) to help deliver modern, secure data solutions for a high-impact federal program. In this role, you will support the development of scalable cloud lakehouse capabilities, data pipelines, access-control frameworks, applications, and APIs that enable federal agencies to integrate, analyze, and share mission-critical data. Candidates must be able to pass a background check equivalent to a Federal Public Trust clearance. This is a remote position, with a requirement to come into the office quarterly for in-person meetings.
Key Responsibilities
• Design, build, and maintain data pipelines on Databricks using Python and PySpark, including schema normalization, metadata capture, data quality checks, and lineage tracking.
• Develop and improve canonical/unified data model (UDM) structures, including star schema, fact, and dimension table design, to support repeatable onboarding of new agency datasets.
• Build and integrate Retrieval-Augmented Generation (RAG) and vector search capabilities to support AI-assisted contract search and natural language querying.
• Contribute to Databricks Asset Bundles (DAB) and CI/CD pipelines to improve deployment consistency, testing, and release velocity.
• Develop Python-based integration layers connecting the OPA/Rego policy engine to Databricks and Spark, enabling dynamic enforcement of access controls and data masking at query time.
• Support Delta Sharing configurations, text extraction/OCR workflows, and Databricks Apps as the team expands into these areas.
• Package, document, and harden pipeline and data model patterns into reusable, well-documented reference implementations that other agencies and teams can adopt independently.
• Participate in weekly standups and contribute to monthly status reporting on task progress and milestones.
Databricks Focus Areas for This Role
Candidates should have direct, hands-on experience in one or more of the following, and be comfortable ramping quickly across the rest:
• Complex pipeline development and improvement on Databricks/PySpark
• Star schema and unified data model (UDM) design and refactoring
• RAG / vector search implementation for search and natural language querying use cases
• Fact and dimension table design for analytic workloads
• Databricks Asset Bundles (DAB) and CI/CD best practices
• Delta Sharing, text extraction/OCR, and Databricks Apps (secondary priority areas)
Required Qualifications
• 3–5+ years of hands-on data engineering experience, including at least 1–2 years working directly in Databricks.
• Strong proficiency in Python and PySpark for building and troubleshooting data pipelines.
• Working knowledge of dimensional data modeling concepts (star schema, fact/dimension tables) and willingness to grow this into a core strength.
• Experience with SQL and working across structured and semi-structured data sources.
• Familiarity with CI/CD concepts and version-controlled deployment workflows (Git-based).
• Comfortable working in a fast-paced, collaborative team environment and picking up new tools quickly.
• Ability to obtain/maintain a Federal Public Trust clearance.
Preferred Qualifications
• Exposure to Databricks Asset Bundles (DAB) or similar infrastructure-as-code deployment tooling.
• Exposure to vector search, embeddings, or RAG-style architectures — specifically vector database/embedding tooling (e.g., Databricks Vector Search, pgvector, FAISS, or Chroma) and embedding or generation model integration (e.g., Databricks Model Serving, Azure OpenAI, Bedrock).
• Familiarity with Open Policy Agent (OPA) / Rego or other policy-as-code frameworks.
• Experience with a general-purpose backend language and a modern frontend framework for adjacent API or UI work.
• Prior experience on a federal contract or in a regulated data environment.
• Databricks certification(s) (Data Engineer Associate/Professional, or Generative AI Engineer Associate).
• Familiarity with Unity Catalog governance features (fine-grained access control, row/column-level security, data lineage).
• Familiarity with federal compliance frameworks (e.g., NIST 800-53, FISMA, ATO processes) or experience handling CUI/PII.
• An active Public Trust (or higher) clearance or investigation already in process.
• Experience with data quality/testing frameworks
The salary range provided represents the estimated compensation for new hires in this position, applicable across all locations. Actual offers may vary based on factors such as the candidate's skills, qualifications, experience, and market conditions. Full Tilt Data complements its base salary offering with a competitive package that includes health benefits, discretionary bonuses, and reimbursement for professional development and training.
Full Tilt Data provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training.
- Pay Range
- $125,000.00 None to $150,000.00 None