Senior Software Engineer I

Hace 23 horas

Mexico City PowerToFly Jornada completa

Job Description

Job Description

As a Senior Software Engineer, Customer Files Orchestration, you will design, build, scale, and operate the service that makes customer documents usable by CoCounsel Legal agents. You will work on the end-to-end processing of customer files, including upload, storage, extraction, embedding, indexing, and orchestration across distributed systems.

You will focus on developing reliable, scalable, and observable Python services on AWS and Kubernetes. The role requires hands‑on expertise in distributed systems, asynchronous processing, queuing, flow control, consistency models, API design, and operational ownership. You will collaborate with engineering teams responsible for storage, extraction, indexing, connectors, and front‑end experiences to deliver dependable customer file workflows.

Key Responsibilities:

  • Own the Orchestration Service End-to-End: Design, implement, test, deploy, scale, and operate meaningful components of the customer files orchestration service.
  • Design Distributed Systems Solutions: Define and document technical designs, including decisions related to consistency, delivery guarantees, crash recovery, asynchronous service boundaries, and durable state management.
  • Build Queuing and Flow‑Control Capabilities: Develop fair queuing, rate limiting, backpressure, circuit breaking, dead‑letter handling, and autoscaling mechanisms to support high‑volume, multi‑tenant workloads.
  • Develop Reliable File Processing Workflows: Support the orchestration of file upload, storage, extraction, embedding, and indexing workflows across multiple services and processing stages.
  • Design APIs as Products: Build and maintain versioned APIs with documented behavior, backward compatibility, clear ownership, and deliberate change management.
  • Ensure Correctness at Scale: Anticipate concurrency, load, partial‑failure, retry, and long‑running processing scenarios, and design systems that prevent recurring classes of defects.
  • Operate What You Build: Participate in primary on‑call rotations, lead incident response activities, and drive post‑mortem actions that prevent repeated operational issues.
  • Improve Observability: Instrument and trace services to enable timely diagnosis of file‑processing workflows, downstream notifications, and customer‑impacting issues.
  • Reduce Operational Toil: Automate recurring manual work and convert undocumented operational processes into reliable, repeatable runbooks or system improvements.
  • Collaborate Across Engineering Teams: Partner with storage, extraction, indexing, connectors, and front‑end teams to ensure integrated and effective file‑processing experiences.
  • Mentor and Support Engineering Excellence: Contribute through code reviews, design reviews, technical guidance, and knowledge sharing across the team.
  • Communicate Technical Trade‑Offs: Clearly explain system design decisions, delivery risks, throughput considerations, and the cost of correctness to engineering and product partners.

Required Qualifications:

  • Bachelor's or master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • At least 3 years of experience building production software and owning significant components within customer‑facing systems.
  • Extensive experience with Python and AWS in production environments.
  • Experience developing and debugging concurrent Python applications, including asyncio‑based workflows.
  • Understanding of Python concurrency concepts, including the Global Interpreter Lock, event loops, semaphores, worker pools, blocking operations, and background task behavior.
  • Experience running Kubernetes workloads in production environments, including event‑driven autoscaling with KEDA, scaler tuning, resource requests, and production troubleshooting.
  • Experience building large‑scale orchestration services or data pipelines that process high volumes of work across multiple service boundaries.
  • Practical understanding of eventual consistency, strongly consistent reads, durable state, retries, partial failures, and long‑running processing workflows.
  • Experience with delivery guarantees, including at‑least‑once processing, exactly‑once considerations, and idempotent processing.
  • Hands‑on experience with queues, fair queuing, rate limiting, backpressure, multi‑tenant fairness, and noisy‑neighbor isolation.
  • Experience designing and maintaining APIs with versioned contracts, backward compatibility, and clear ownership.
  • Demonstrated operational ownership, including on‑call support, incident response, root‑cause analysis, and implementation of durable fixes.
  • Strong engineering fundamentals, including testing, code review, observability, CI/CD, and infrastructure as code