We're looking for an engineer who combines backend engineering ability with LLM/Prompt engineering ability — someone who both develops real-time in-class backend services and owns the design, debugging, and continuous optimization of teaching prompts and conversational ability, so the AI teacher speaks better, more stably, and more controllably.
Responsibilities
- Design and develop real-time in-class teaching backend services, ensuring the conversation pipeline is stable, low-latency, and highly available.
- Own the design, debugging, and version management of teaching prompts (lesson types / question types / teacher scripts / multi-turn dialogue strategy), supporting real classroom teaching outcomes.
- Design and implement engineering capabilities for LLM invocation: multi-model routing, streaming output, context management, Function/Tool Calling, timeout/retry/fallback (degradation), cost and latency optimization.
- Build an evaluation and regression system for prompt and dialogue effectiveness (metric design, bad-case analysis, A/B testing, regression sets), driving measurable, continuous improvement of teaching outcomes.
- Handle LLM stability and safety issues: hallucination suppression, output constraints, content safety/compliance filtering.
- Collaborate with instructional design, product, and algorithm teams to accurately translate teaching intent into controllable, reusable prompts and service capabilities.
Requirements
Backend Engineering Ability
- Bachelor's degree or above, Computer Science-related major, 5+ years of backend development experience.
- Proficient in Java (Spring ecosystem) or equivalent backend ability; familiar with microservices, high concurrency, and interface design; has JVM tuning experience.
- Familiar with real-time/streaming pipeline development (SSE/WebSocket, reactive programming such as Reactor or WebFlux); understands low-latency and backpressure handling.
- Familiar with caching (Redis), message queues, and database design; capable of production issue diagnosis and performance optimization.
AI/Prompt Engineering Ability (Core)
- Solid hands-on Prompt engineering experience: proficient in system/role prompt design, few-shot, Chain-of-Thought (CoT), structured output (JSON Schema / constrained decoding), context window management and trimming, multi-turn dialogue state control.
- LLM application engineering ability: familiar with mainstream LLM APIs and orchestration, skilled in streaming token processing, Function/Tool Calling, multi-model routing/switching, token and cost control, time-to-first-token optimization, timeout/retry/fallback.
- Able to iterate systematically based on evaluation: can build evaluation and regression mechanisms (automated evaluation, bad-case attribution, A/B comparison, human annotation feedback loop), using data to drive prompt optimization rather than intuition.
- Understands the capability boundaries and sampling parameters of mainstream LLMs (temperature/top-p, etc.), and the applicable scenarios and trade-offs of embeddings and RAG/fine-tuning.
- Has real hands-on experience handling LLM hallucination, stability, and content safety issues.
Bonus Points
- Familiar with LLM orchestration/application frameworks (Spring AI, LangChain, etc.).
- Experience with conversational, educational/teaching, or multi-agent (orchestration, memory, instruction parsing) systems.
- Experience designing RAG, vector retrieval, or long/short-term memory / context systems.
- Experience building LLM evaluation platforms or prompt versioning/experimentation platforms.
- Familiar with how TTS/ASR voice pipelines coordinate and optimize latency in real-time dialogue scenarios.
- Keeps up with cutting-edge LLM developments and can apply new methods in practice.