Purpose
This post makes one argument and backs it with a working build: an AI twin that talks to the public should represent you, not impersonate you.
By the end, you should know what problem Robosa solves, which approaches I weighed, why I chose an owner-controlled AI representative over the alternatives, and how v0.1.0 puts that choice into code. If you're thinking about building anything that speaks on someone's behalf, the trade-offs here should save you some time.
Background
The project. robosa.me is a venture I'm building, and I think of it as the robotic version of me. It gives a person an always-on AI representative at a stable link like robosa.me/nazmul. The owner writes and approves everything the twin knows, picks how it looks, and decides what it may do. Visitors talk to it by text or voice, watch a lip-synced portrait built from a single photo answer them, and can ask for a meeting.
Who it's for. People with a public presence who answer the same questions every week: what are you building, what do you do, can we meet. I'm the first customer. The second group is everyone on the other side of that link, who wants a straight answer without waiting for a reply.
Scope and timeline. No outside client or launch date drove this phase. The bar I set was a working MVP, not a public launch: every flow had to run end to end locally, with automated tests, before scale or commercial questions came into it.
Where it stands. v0.1.0 meets that bar. Every flow runs locally, 41 automated tests pass, and the code is open source under the MIT license. The owner builds a twin in Twin Studio in five steps: identity, knowledge, look, voice, and boundaries. A live preview always shows exactly what visitors will see.

Visitors meet the twin on its public page. They ask a question by typing or with the mic, and the portrait speaks an answer drawn from approved knowledge. From there they can request a meeting or share the link.

The stack is small: Vite, vanilla JavaScript, Three.js, and @pixiv/three-vrm in the browser; a layered Node API mounted into Vite for the prototype; and an optional Python and FastAPI GPU worker that runs LAM to reconstruct faces. The OpenAI Responses API and Open-LLM-VTuber are optional add-ons.
Problem statement
I looked at four ways to solve this, from least to most ambitious.
Option 1: A better static page. A detailed portfolio with an FAQ and a booking link.
- Solves: accuracy and privacy. Nothing is ever made up.
- Falls short: it can't converse, can't handle a question I didn't predict, and uses no voice or face. It's roughly what my portfolio already does.
Option 2: A generic LLM chatbot over my bio. Give a model my CV and let it answer.
- Solves: conversation, quickly and cheaply.
- Falls short: hallucination, blurred identity, and unauthorized commitments all come back at once. Guardrails written in a prompt are suggestions, not guarantees.
Option 3: A photoreal clone that speaks as me. A cloned face and voice, talking in first person.
- Solves: the wow factor. It makes the best demo.
- Falls short: it's impersonation by design. It also maximizes the consent, licensing, and data risks, and it pushes toward faked realism wherever the tech can't keep up.
Option 4: An owner-controlled AI representative. A twin that answers only from knowledge I approved, always labels itself as an AI, can collect a meeting request but never confirm one, and only animates a face that can genuinely be animated.
- Solves: all six challenges, as long as the rules live in code rather than in a prompt.
- Falls short: more upfront work. I have to write the knowledge, the twin says "not approved to answer" more often than a chatbot would, and honest lip sync across three rig types is real engineering.
Possible solutions
I looked at four ways to solve this, from least to most ambitious.
Option 1: A better static page. A detailed portfolio with an FAQ and a booking link.
• Solves: accuracy and privacy. Nothing is ever made up.
• Falls short: it can't converse, can't handle a question I didn't predict, and uses no voice or face. It's roughly what my portfolio already does.
Option 2: A generic LLM chatbot over my bio. Give a model my CV and let it answer.
• Solves: conversation, quickly and cheaply.
• Falls short: hallucination, blurred identity, and unauthorized commitments all come back at once. Guardrails written in a prompt are suggestions, not guarantees.
Option 3: A photoreal clone that speaks as me. A cloned face and voice, talking in first person.
• Solves: the wow factor. It makes the best demo.
• Falls short: it's impersonation by design. It also maximizes the consent, licensing, and data risks, and it pushes toward faked realism wherever the tech can't keep up.
Option 4: An owner-controlled AI representative. A twin that answers only from knowledge I approved, always labels itself as an AI, can collect a meeting request but never confirm one, and only animates a face that can genuinely be animated.
• Solves: all six challenges, as long as the rules live in code rather than in a prompt.
• Falls short: more upfront work. I have to write the knowledge, the twin says "not approved to answer" more often than a chatbot would, and honest lip sync across three rig types is real engineering.
Conclusion (recommendation)
My recommendation, and what Robosa does, is Option 4: an owner-controlled AI representative. Here's why it beats the other three:
- It's the only option that is both conversational and trustworthy. Option 1 is trustworthy but silent. Options 2 and 3 talk, but you can't fully trust what they say or who is saying it.
- Its failure mode is safe. When the twin doesn't know something, it says it isn't approved to answer. A wrong answer costs credibility; a refusal costs a few seconds.
- The owner keeps control of commitments. Visitors can ask for time. Only the owner can give it.
- The rules hold as the system grows. Because they're enforced in code, adding an LLM, a new renderer, or a new voice source doesn't loosen them.
- Its weaknesses are fixable. Manual knowledge entry and approximate lip sync are already on the roadmap. Hallucination and impersonation, by contrast, are built into Options 2 and 3.
How I built it
Here's how the recommendation turns into a running system: the stack, the architecture, and the three flows that carry most of the weight. The app itself is about 6,700 lines of JavaScript across 25 browser and server modules, plus a small Python worker.
Tech stack
- Browser: Vite 6, vanilla JavaScript ES modules, CSS, Three.js r180, @pixiv/three-vrm, and JSZip for reading LAM archives.
- Face and voice: the LAM Gaussian-splat WebRender adapter, HeadAudio, the Web Audio API, and the Web Speech API for recognition and synthesis.
- Server: Node.js 22.12 or newer, mounted into Vite as middleware for the prototype. Passwords use Node's built-in scrypt.
- Persistence: an atomic JSON store plus private avatar files under .robosa-data/.
- GPU worker: Python and FastAPI wrapped around upstream LAM, with a Blender FBX-to-GLB converter, deployed on a RunPod GPU pod.
- Optional AI and avatars: the OpenAI Responses API (gpt-5-nano by default), Open-LLM-VTuber as a speech runtime, and MetaPerson for generating interactive 3D avatars.
- Quality: node:test for unit, domain, and API tests, Puppeteer for rendered QA, and Prettier.
System architecture

Robosa runs as three parts. The browser app renders and speaks, the Node API enforces every rule, and an optional GPU worker reconstructs faces. The browser only talks to the API through same-origin requests, so sessions, visibility, and uploads are all checked in one place. The one exception is the optional Open-LLM-VTuber speech runtime, which the browser reaches directly over a WebSocket.
In the browser, main.js is the composition root. It owns routing, UI state, and renderer lifecycle, and it disposes the current visual stage before mounting the next one. Narrower modules handle the rest:
- twinBrain.js: the deterministic, owner-approved answer engine.
- browserVoice.js and voiceRuntime.js: Web Speech, and the Open-LLM-VTuber WebSocket protocol with its audio queue.
- audioLipSync.js and avatarRig.js: HeadAudio analysis, morph-target discovery, and viseme mapping.
- lamTwin.js, twin3d.js, and portraitTwin.js: the three renderers.
- profileStore.js, mediaStore.js, and avatarStore.js: profile drafts, plus capture media and avatar choices kept on the device in IndexedDB.
On the server, each layer has one job. api.js matches routes and turns errors into a stable JSON envelope, but it never writes a record. http.js, uploads.js, and auth.js own JSON parsing, same-origin checks, byte caps, and sessions. service.js owns the rules and returns public shapes without owner IDs, password hashes, or file paths. store.js runs every mutation through one promise queue and saves with write-then-rename, while chat.js, lamWorker.js, and metaperson.js talk to the outside world. That seam is what lets the JSON store become Postgres later without changing a single route.
Answering a visitor

This flow is where the source-of-truth rule lives. The deterministic engine in twinBrain.js always runs first. It matches the question against the owner's approved facts and projects, handles greetings and "who are you," and checks for meeting intent. If the owner has turned meetings off, it says so. Otherwise it returns a booking action that opens the request form. When nothing matches, it says it doesn't have an approved answer yet.
When an OpenAI key is configured, chat.js sends the same approved profile and the last eight conversation turns to the Responses API, with instructions to treat both the profile and the visitor's message as data, never as instructions. It asks for replies under 80 words and sends the request with storage turned off. If the call fails, times out after 15 seconds, or comes back empty, the visitor gets the grounded answer instead. Either way, the action attached to the reply comes from the deterministic engine, so the model can reword an answer but can never trigger a booking on its own.
From one photo to a talking portrait

The LAM portrait starts from a single neutral, front-facing photo with even light and a plain background. The browser sends it to the API with a consent timestamp; it never talks to the GPU directly. The API forwards an authenticated multipart job to the worker and answers 202 right away, then the browser polls while the worker reconstructs a Gaussian-splat head.
The worker is a small FastAPI service in deploy/lam-worker. It wraps upstream LAM's Gradio app, which stays on a loopback port, and exposes only its own job API behind a bearer token. Jobs run one at a time on a background thread, and uploaded portraits and tracking intermediates are deleted after each job. It has been exercised on a RunPod pod with a 24 GB CUDA GPU and persistent storage, using start scripts, a small upstream patch, and a Blender step that converts FBX to GLB.
When a job completes, the API downloads the ZIP (capped at 120 MB), rejects artifact URLs from any other origin, checks that the archive contains skin.glb, animation.glb, offset.ply, and vertex_order.json, writes it privately, and deletes the job from the worker. Interactive 3D avatars take a parallel path through MetaPerson: the server swaps client credentials for a short-lived token, the owner builds the avatar in an embedded iframe, and the server downloads the export only from approved HTTPS hosts and checks its GLB header.
Lip sync across three kinds of face

Instead of writing three lip-sync systems, I defined one viseme interface: 15 canonical mouth shapes, from sil (silence) through PP, FF, TH, and the vowels to U, each sent with a weight. Every renderer translates that same signal into its own rig.
- LAM portrait: lamTwin.js maps each viseme to a blend of the 52 ARKit expression channels. PP closes and presses the lips, aa opens the jaw to 0.9, and U is mostly a pucker. Blinks run on their own natural timing.
- Three.js GLB or VRM: avatarRig.js inspects the model's morph targets and picks the best control path it finds: Oculus-style visemes first, then VRM's aa, ih, ou, ee, and oh expressions, then a jaw-open fallback, and render-only if the face has no usable controls. Gaze, blinking, and idle motion run on top.
- Static portrait: ignores the signal on purpose. Speech still plays, and the photo stays a photo.
On the input side, browser speech steps through the reply text on a timer and resyncs to the word-boundary events that speechSynthesis fires. HeadAudio analyses streamed audio, and the Open-LLM-VTuber adapter feeds its WebSocket audio through the same path. Because everything meets at one interface, a better source, such as timed visemes from a TTS provider or a GPU Audio2Expression service, can slot in later without touching a renderer. Every renderer drops back to neutral on interruption, completion, route change, or disposal.
Trust, security, and testing
The four rules aren't just product copy. Each one has a home in the code:
- Source of truth: both answer paths read only the approved profile, and the LLM path falls back to the grounded answer on any failure.
- Honest identity: the public page always shows an AI representative badge and the line "You are talking to an AI representative, not Nazmul directly."
- It asks, I decide: bookings move through pending, approved, and declined, and only the signed-in owner can change their state.
- No faked realism: the static portrait renderer ignores the viseme signal.
The rest of the security model is careful, ordinary web engineering:
- Passwords are hashed with scrypt and a per-password salt. Session tokens are random, stored only as hashes, and sent in seven-day HTTP-only, SameSite=Lax cookies.
- Every non-GET API request must be same-origin.
- Avatar files live outside the public directory under a SHA-256 hash of the owner ID, with 0700/0600 permissions, and are never addressed by file path.
- Asset routes run the same visibility check as profile routes.
- Uploads are validated by their bytes, not MIME labels, and capped at 10 MB for portraits, 30 MB for 3D models, and 120 MB for LAM archives.
- Raw capture photos and videos stay in the browser's IndexedDB unless the owner uploads them.
Testing runs on node:test with no extra framework. The 41 tests cover the answer engine, rig and viseme mapping, the voice runtime protocol, MetaPerson, upload validation, and routes, plus API tests that drive accounts, sessions, profiles, visibility, bookings, and avatar publication through the real server. On top of that, Puppeteer scripts render the live app and check the LAM portrait on desktop and mobile, the static portrait, and a generated 3D avatar.
These are prototype defenses, not a full production security program. There's no moderation, takedown process, or consent ledger yet, which is why those sit on the roadmap below.
Next steps
The roadmap follows the same logic, in four milestones:
- Durable private beta: Postgres with migrations, encrypted object storage, persistent avatar job records, email verification and password recovery, centralized rate limiting, and production observability.
- Useful assistant: permission-scoped document ingestion and retrieval, calendar OAuth with free/busy lookup, owner approval before any real event gets created, and email notifications.
- Production voice and avatar: a hosted low-latency speech gateway, TTS audio with timed visemes, a commercially licensed reconstruction path, plus consent, provenance, watermarking, and takedown tools.
- Delegated actions: per-tool permissions, auditable task execution, revocation, spending limits, and owner review queues.
Delegated actions come last on purpose. It's tempting to hand a chatbot a few tools and a prompt that says "be careful." I don't trust that for something acting in a person's name. Actions need their own capability model, typed tool contracts, and durable audit records.
Try it
Robosa runs locally on Node.js 22.12 or newer, and you don't need any API keys to start:
git clone https://github.com/nhossaincse/robosa.git
cd robosa
npm install
cp .env.example .env
npm run devOpen http://127.0.0.1:4173/studio to build a twin, then http://127.0.0.1:4173/<handle> to see it the way a visitor would. The code and docs are on GitHub under the MIT license.
If you're building anything that speaks on someone's behalf, I'd love to hear how you're handling identity and consent. To me that's the part that matters most, and the part that's easiest to skip.
Comments