Role Summary
We are looking for a Python Developer — Data Engineering to build and operate the ingestion and transformation backbone of an AI-driven contact centre analytics platform.
You will own pipelines that pull interaction data and recordings from NICE CXone, process them through transcription and PII redaction, normalize them into a relational model, and populate vector collections supporting AI workloads.
Key Responsibilities
- Build and maintain Python ETL pipelines ingesting call recordings and interaction metadata from NICE CXone.
- Work with the Storage Export API and Data Extraction API, including tenant concurrency and retention-window considerations.
- Build transformation layers to normalize raw payloads into the platform's relational schema.
- Integrate Vertex AI Speech-to-Text, including channel configuration, diarisation and confidence scoring.
- Implement PII detection and de-identification using Google Cloud Sensitive Data Protection.
- Build embedding pipelines writing to Qdrant collections for conversations, utterances and agents.
- Implement hybrid dense/sparse retrieval and tune HNSW parameters for recall and latency.
- Develop Airflow / Cloud Composer DAGs with retry, backfill, idempotency and alerting.
- Build and maintain internal APIs using FastAPI or equivalent.
- Implement pipeline observability covering throughput, latency, failure rates and data freshness.
- Participate in production support and incident resolution.
- Write unit and integration tests, participate in code reviews and maintain technical documentation.
Must-Have Skills
- Strong Python 3.10+ production experience.
- pandas, SQL Alchemy, async I/O, packaging and pytest.
- Experience building batch and incremental ETL pipelines at scale.
- Advanced SQL.
- Hands-on experience with BigQuery and at least one OLTP engine such as Cloud SQL, PostgreSQL or SQL Server.
- Airflow or Cloud Composer experience.
- Working knowledge of GCP — Cloud Run, Cloud Storage, IAM, service accounts and Secret Manager.
- REST API integration experience, including OAuth 2.0, pagination, rate limiting and retry/backoff strategies.
- Docker/containerisation.
- Git-based development and CI/CD workflows.
Preferred / Nice to Have
- Vector database experience — Qdrant strongly preferred; Pinecone, Weaviate or pgvector acceptable.
- Experience with embedding models and hybrid dense/sparse retrieval.
- Speech-to-text or audio-processing experience.
- Exposure to NICE CXone, Genesys, Five9 or Amazon Connect.
- Experience with DLP, tokenisation or masking in regulated environments.
- Financial services or contact-centre domain experience.
- Terraform / Infrastructure-as-Code experience.