AI Enginner
Software Engineer, AI Agent Platform
Company: Happiness with Apple Technology Co., Limited
Department: IT
Location: Hong Kong
Employment Type: Full-time
Work Mode: On-site
Salary Range: HKD 25,000-35,000 per month
About the Company
Founded in Hong Kong in 2018, Happiness with Apple Technology Co., Limited is committed to helping customers succeed in global markets. We are building a one-stop overseas service platform for cross-border businesses, providing localized overseas operations, traffic growth, and payment collection solutions.
Our business now covers dozens of countries and regions, with eight branch offices across Mainland China, Hong Kong, Singapore, Malaysia, India, and other markets.
About the Role
You will help build and operate our production agent runtime. You will work alongside senior engineers on prompt orchestration, function calling, MCP infrastructure, memory systems, retrieval pipelines, and the safety and observability layer that makes agents production-ready.
This role is a good fit if you have a solid backend foundation and have shipped at least one LLM-powered feature to production, even if you have not owned a full agent system end to end yet. We will pair you with senior engineers and expect you to grow into deeper technical ownership over time.
What You’ll Do
Agent Runtime & Tooling
Build and improve components of our agent runtime, including tool calling, workflow execution, state management, and failure recovery.
Implement prompt rendering, routing, and orchestration patterns for tool chains across multiple workflows and models.
Contribute to our internal MCP infrastructure and build standardized connectors for APIs, databases, and SaaS tools.
Help build reusable SDKs and engineering patterns that other teams use to ship their own agents.
Safety & Reliability
Implement safety controls such as input/output validation, prompt injection defenses, data masking, and audit logging.
Add retry, timeout, fallback, and graceful degradation behavior to agent execution flows.
Investigate and fix reliability issues caused by unstable model APIs, slow tools, partial failures, and long-running tasks.
Help define practical guardrails for privacy, compliance, and safe tool execution.
Performance & Observability
Profile and tune retrieval, reranking, prompt construction, and tool-call paths to improve latency and cost.
Build observability into the runtime, including traces, metrics, evaluations, and debugging tools.
Help monitor agent quality, cost, reliability, and failure modes in production.
Turn agent prototypes from notebooks or experiments into maintainable production code.
What We’re Looking For
Core Engineering
3-5 years of backend, platform, or infrastructure engineering experience.
Fluent in Python, Java, or Node.js. We lean Python.
Production experience with relational databases, Redis, and async messaging systems such as Kafka, SQS, RabbitMQ, or similar.
Comfortable building services with FastAPI, Spring Boot, or comparable backend frameworks.
Familiar with distributed systems basics, including concurrency, retries, idempotency, and graceful failure handling.
Hands-on experience with Docker, Git-based workflows, and CI/CD. Kubernetes exposure is a plus.
Retrieval & Data Systems
Experience with at least one vector database such as pgvector, Pinecone, Chroma, Milvus, Weaviate, or similar.
Familiar with embeddings, hybrid search, metadata filtering, and basic RAG pipeline patterns.
You do not need to be an HNSW or IVF expert, but you should know what the main tuning knobs are.
Agent & LLM Systems
Hands-on experience with at least one LLM or agent framework such as LangChain, LangGraph, LlamaIndex, AutoGen, or similar.
You have shipped at least one LLM-powered feature or RAG system to production.
You have operational stories from running LLM systems: eval gaps, prompt regressions, cost surprises, latency issues, or tool-call failures.
Working knowledge of tool/function calling flows and basic agent state management.
Deep MCP experience is not required, but you should be excited to learn it quickly.
Nice to Have
Experience with workflow engines such as Airflow, Temporal, Dagster, or similar.
A GitHub portfolio with backend, AI, infrastructure, or developer tooling projects.
Exposure to evaluation systems, hallucination detection, or LLM observability tooling.
Familiarity with private LLM deployment using vLLM, TGI, Ollama, or similar systems.
Awareness of AI safety, privacy, compliance, or audit logging considerations in production systems.
