About the job
We are looking for a Senior Software Engineer to work on the indexing and data processing layer of a novel search engine tailored for agentic AI consumption.
In this role, you will focus on building systems that ingest, process, and organise massive volumes of data into efficient, queryable structures. You will work primarily on offline and nearline pipelines, ensuring that data is fresh, complete, and efficiently accessible by downstream retrieval systems. You will operate in an environment where throughput, scalability, and correctness are critical; designing systems capable of handling tens of gigabytes per second across continuously evolving datasets.
Responsibilities
Design, implement, and operate large-scale indexing systems and data pipelines that sit at the core of our search infrastructure
Develop and optimise indexing strategies balancing performance, freshness, and resource efficiency
Work on storage formats, compaction strategies, and update mechanisms to keep data accessible and current
Ensure reliability and predictability of pipelines under high-throughput conditions
Build well-tested components with clear responsibilities and interaction contracts, while remaining flexible as the system evolves
Define and implement observability primitives, including structured logs, metrics, and data quality signals across offline and nearline pipelines
Monitor throughput, resource usage, and cost, and drive optimisations when business needs require it
Collaborate with runtime and ML teams to ensure indexing outputs meet retrieval and ranking requirements
Enable safe experimentation on indexing strategies and data processing logic through controlled rollouts and clearly defined quality signals
Qualifications
Minimum
5+ years of experience building production backend or data infrastructure systems
Strong Go experience (C++/Rust is a plus)
Experience with large-scale data processing systems (10+ GiB/sec throughput, petabyte-scale datasets, etc.)
Experience building or operating databases, storage systems, data planes, or indexing pipelines
Strong understanding of distributed systems, fault tolerance, consistency, and scalability
Experience running production systems and handling operational incidents
Systems-thinking mindset and ability to reason about end-to-end data flows
Preferred
Distributed data processing frameworks such as Spark, Flink, MapReduce, or Beam
Content systems including, scraping, proxying, or anti-bot infrastructure
Ad tech, social networks, or other large-scale content platforms
DBMS internals (open source or SaaS) and cloud infrastructure
Open-source contributions or active involvement in the engineering community
Competitive programming or CTF participation (ICPC, IOI, or similar)
SHAD or similar advanced technical programmes
Conference talks or technical publications