Overview
Airport touchscreens are slow, crowded, and not especially hygienic. Aria replaces them with a voice-first assistant that feels like talking to a knowledgeable airport employee: walk up, get recognized, and ask about your flight.
How it works
- Sign-in. Passengers authenticate with their face (DeepFace Facenet512 embeddings matched against a MongoDB Atlas vector index) or with a QR code.
- Conversation. Speech streams over WebSockets to a FastAPI backend, where a LangGraph agent running on Vultr Serverless Inference (Kimi-K2-Instruct) decides which tools to call: flight status, gates and boarding times, weather at origin and destination, airport maps, and rebooking.
- Knowledge. Policy questions about baggage and security are answered with retrieval over an airport knowledge base.
- Voice and presence. Replies stream back as ElevenLabs speech with character-level alignment, and a 3D particle sphere reacts to the audio in real time.
What I built
I built the frontend and collaborated on the agent's capabilities. The centerpiece is the assistant itself: a Three.js sphere of 20,000+ particles deformed by a Web Audio FFT of Aria's voice, with bloom post-processing and a small animation state machine (drag → momentum → auto-rotation). Animation state lives in refs rather than React state, which is what keeps it at 60 fps without re-rendering the app.
Challenges
Keeping audio, text, and animation in sync over a streaming WebSocket connection took most of the effort, along with smoothing the FFT so the sphere breathes instead of jittering. We also migrated the backend architecture mid-hackathon and made tool failures degrade gracefully instead of crashing the conversation.