Jobgether
Accountabilities:: Design, build, and maintain data pipelines using BigQuery, Dataflow, Cloud Composer, Apache Airflow, or comparable technologies to ingest, transform, and serve data for GenAI use cases. Develop and maintain ELT/ETL processes for batch ingestion from client source systems such as student information systems, ERP platforms, and casework systems into cloud analytics and RAG data stores. Monitor pipeline reliability, performance, and cost while optimizing cloud data processing as engagements scale. Classify and tag sensitive or regulated client data using data governance and loss-prevention tools such as Dataplex and Cloud DLP. Define and enforce appropriate data access controls and governance policies for higher education and public-sector information. Assess data readiness for AI applications and proactively identify gaps in data quality, completeness, access, or governance. Design data models and schemas that support both immediate client requirements and reusable patterns across future engagements. Partner with client technical teams to understand source-system constraints and determine effective data extraction and integration approaches. Maintain clear technical documentation covering data flows, schemas, classification decisions, and governance practices for reuse and audit purposes. Execute delivery work from a scoped backlog, provide technical estimates, and identify data-related risks that could affect timelines or outcomes. Partner with cloud and AI engineering teams to ensure pipelines provide the structure, quality, and freshness required by agentic AI and retrieval-augmented generation solutions. Conduct data quality reviews and testing to maintain reliable standards across projects. Collaborate with infrastructure, security, delivery, and engagement teams on data-related scope, risks, timelines, and technical decisions. Identify opportunities to convert one-off client solutions into scalable and repeatable technical assets. Requirements Bachelor’s degree in a related field or equivalent professional experience. Hands-on experience building data pipelines and ELT/ETL processes for analytics or GenAI use cases on a major cloud data platform such as BigQuery, Snowflake, Redshift, or Synapse. Direct experience with BigQuery and Dataflow, Cloud Composer, or Apache Airflow is strongly preferred. Strong GenAI or AI-adjacent data engineering experience on other platforms may be considered for candidates able to quickly develop Google Cloud expertise. Experience with data governance and classification technologies such as Google Cloud Dataplex and Cloud DLP, or comparable tools including Collibra, Alation, or AWS Macie. Working knowledge of cloud storage, messaging, and access-control concepts such as Cloud Storage/S3, Pub/Sub, SNS/SQS, and IAM. Google Cloud Professional Data Engineer certification is preferred. Strong SQL skills, including complex queries, stored procedures, user-defined functions, and performance tuning. Strong Python skills for pipeline development and data transformation. Proficiency in ETL design and development using at least one relevant tool, such as SSIS. Basic knowledge of C#/.NET for SSIS scripting. Solid understanding of relational database management systems, database design, and data modeling. Familiarity with data privacy and compliance requirements relevant to education or public-sector data, including FERPA and state privacy laws. Experience integrating data from legacy or third-party systems such as SIS, ERP, or casework platforms. Strong understanding of software development lifecycle practices, with Agile or iterative delivery experience preferred. Familiarity with Jira or TFS is preferred; Netezza and PostgreSQL experience is helpful. Strong analytical, troubleshooting, and problem-solving abilities, with sound technical judgment around data quality, access, governance, and reliability. Ability to communicate technical data considerations clearly to both technical and non-technical stakeholders. Ability to work effectively from a scoped backlog and translate technical requirements into production-ready data solutions. Strong ownership, professionalism, adaptability, and resilience in a rapidly changing environment. Ability to maintain a high level of confidentiality when working with sensitive information. Comfortable working independently as part of a distributed virtual team while remaining productive and engaged. Strong written and verbal communication skills. Ability to obtain a security clearance. Benefits Salary range of $115,000–$135,000, based on experience. Full-time remote work arrangement. Medical, dental, and vision coverage. Health Savings Account (HSA) and Flexible Spending Account (FSA) options. Generous earned time off. 401(k) and student loan repayment benefits. Life insurance and AD&D insurance. Short- and long-term disability coverage. Employee Assistance Program. Employee stock purchase program. Tuition reimbursement. Performance-based incentive pay. Robust wellness program. Opportunity to work at the intersection of data engineering, cloud technology, and GenAI. Exposure to complex higher education and public-sector data environments. Collaboration with cloud, AI, infrastructure, and delivery engineering teams. Opportunities to develop reusable data solutions and contribute to evolving AI-enabled products. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1