Meet Pete. He's chill, he's helpful, and he lives in the corner of Wave Studio.
An animated, voice-capable avatar that chimes in with a useful tip or menu setting when you linger on a control, answers questions right from his bubble instead of the chat menu, summarizes whatever you're looking at, and can be driven by your backend over a simple JSON API.
Personality profile
Wave Studio — mock workspace
Hover or focus any control for ~2.4 s and Pete will chime in once per topicDirector console
Everything below goes through the same VLSAvatar.command() entry point that a backend, an LLM tool call, a postMessage from an iframe, or a server-sent-event stream would use. Try it, then wire the JSON schema into your agent.
Say something
Moods
Gestures
Personality dials
Personality mode
Reactions (app events Pete responds to)
Guided tours
Small talk
JSON command (what your AI sends)
Event log
Drop-in and API reference
One script, no dependencies. Works inside the hosted studio, the desktop packages and the VLS Info Mode extension. Voice uses the device's SpeechSynthesis by default and switches to your hosted narration endpoint when you give it one.
Install
<script src="vls-avatar.js"></script>
<script src="wave-studio-tips.js"></script>
<script>
VLSAvatar.init({
name: 'Pete',
tips: WAVE_STUDIO_TIPS, // context → tips (menu path + manual link)
recital: WAVE_STUDIO_FIFTY, // curated tip list the user ticks on and off in settings
knowledge: WAVE_STUDIO_KNOWLEDGE, // offline Q&A fallback
askEndpoint: '/api/agent/ask', // hosted FastAPI → Perplexity Agent API
narrationEndpoint: '/api/narration', // optional: returns audio/mpeg (gpt-4o-mini-tts)
narrationVoice: 'marin',
manualSearchUrl: './manual/?q=',
humor: 0.35, chimeIn: true, chimeGapMs: 45000,
});
</script>
On desktop packages, skip askEndpoint and narrationEndpoint. Pete still answers from the local tip and knowledge packs and uses device speech, matching the offline boundary in the 0.33 builds.
Context tips
Tag any control with data-vls-context="faders". When the user hovers or focuses it for tipDelayMs (default 2.4 s), Pete delivers one unseen tip for that context, then stays quiet for chimeGapMs. Tips can be strings or objects. level: 'pro' tips are held back until the beginner tips for that context have been seen:
VLSAvatar.registerTips({
faders: [
{ text: 'Hold the fader and roll the wheel for fine steps.',
setting: 'Fader → ⋯ → Wheel increment',
manualUrl: './manual/#wheel',
keywords: ['wheel','fine'] }
]
});
VLSAvatar.setContext('faders'); // or let data-vls-context do it
Guided tours
A tour is an ordered list of steps. Each step can spotlight a control (dims the page, pulses an orange ring, scrolls it into view), set the context, show a menu path and link the manual. Users step with Next/Back, arrow keys, or Esc to leave. Completed tours are remembered so first-timers get offered "Show me around" once.
VLSAvatar.registerTours({
'first-look': { title: 'First look', steps: [
{ selector: '[data-tour="stage"]', context: 'stage', text: 'This is the stage…', setting: 'Stage → click', manualUrl: './manual/#selection' },
{ text: 'That\'s the loop. Hover anything for a tip.', mood: 'happy', gesture: 'thumbs' },
]}
});
VLSAvatar.tour('first-look'); // or ask him: "give me a tour"
VLSAvatar.point('[data-tour="faders"]', 'These four.'); // one-off spotlight
Personality
Three modes: chill (default), quiet (no chime-ins, no idle chatter, jokes nearly off) and hype (more jokes, quicker chime-ins). Small talk is handled locally: greetings, thanks, "who are you", "what can you do", jokes. App events can trigger reactions: VLSAvatar.react('saved' | 'exported' | 'recorded' | 'first-cue' | 'error' | 'live-input-on' | 'live-input-off' | 'collision' | 'idle5m'). Memory (visits, tips seen, tours done, mute, mode) lives in localStorage; forget() clears it.
VLSAvatar.setMode('quiet');
VLSAvatar.react('saved'); // "Saved. Future you says thanks."
VLSAvatar.setPersonality({ humor: 0.6, chimeGapMs: 30000 });
Ask & summarize
Clicking Pete opens his bubble with an ask box, so users don't need to leave for the chat menu. ask() tries your handler or endpoint first, falls back to local knowledge, and finally offers a manual search link. summarize() takes a CSS selector, an element or raw text.
VLSAvatar.ask('How do I set up grandMA3 unicast?');
VLSAvatar.summarize('#faders-panel', { title: 'Faders' });
// Custom handler (bring-your-own provider, same as Settings → Assistant)
VLSAvatar.setAskHandler(async (question, { context, history }) => {
const r = await myProvider.chat(question, context);
return { answer: r.text, manualUrl: r.anchor, sources: r.sources, mood: 'happy', gesture: 'nod' };
});
Expected endpoint contract for askEndpoint: POST {question, context, previous_response_id, page} → {answer, manualUrl?, sources?, response_id?, mood?, gesture?}. That maps directly onto the existing /api/agent/ask wrapper.
AI control
Three transports, one schema. Pick whichever your backend already speaks.
| Transport | How |
|---|---|
| Direct | VLSAvatar.command({ action:'say', text:'…' }) — from page scripts or an LLM tool-call handler. |
| postMessage | window.postMessage({ type:'vls-avatar', command:{…} }, '*') — from iframes, companion windows or the Info Mode extension. |
| Server-sent events | VLSAvatar.connectStream('/api/avatar/stream') — backend pushes one JSON command per data: line. See server_example.py. |
Give the LLM the schema below as a tool definition and let it decide when to chime in. A good system rule: "Only send a tip when the user has lingered on a control for more than two seconds, never more than once per minute, keep it under 30 words, and include the menu path."
Command schema
| action | fields | effect |
|---|---|---|
say | text, mood?, gesture?, speak?, sticky?, actions?[{label,href}], sources? | Speak a line with optional buttons and sources |
tip | text, setting?, manualUrl? — or context | Deliver a tip (bulb badge, point gesture) |
ask | question | Run the ask pipeline |
summarize | selector | text, title? | Summarize page content |
mood | chill · happy · thinking · laugh · tip · sleepy | Set facial expression |
gesture | wave · point · shrug · nod · thumbs · bounce · scratch-nose · scratch-head · yawn · stretch · eye-roll · fart · sneeze · look-around · adjust-glasses · adjust-cap · tap-foot · wiggle · facepalm · sweat | Play a one-shot gesture (one at a time; later calls replace the running one) |
moveTo / resetPosition | x, y (viewport px of his top-left) | Place Pete programmatically; users can also hold or drag him anywhere (position is remembered) |
walk / walkHome | selector — or x, y | Walk over to an element (he does this automatically for tour steps and point, then walks back home) |
fidget | gesture? | Play a random idle fidget (never fart) or a named one |
context | context, chime? | Tell Pete where the user is |
mute / stop | value | Voice control |
show / hide / dismiss | — | Visibility |
tour | tour (id) · step · end | Start, step or end a guided tour |
point | selector, text?, durationMs?, sticky? | Spotlight a control with an optional line |
react | event, text? | Respond to an app event (saved, exported, error…) |
mode | chill · quiet · hype | Switch personality mode |
name / tips | value / tips map | Rename, add tips at runtime |
Events
Subscribe with VLSAvatar.on(type, fn) or listen for the vls-avatar CustomEvent on document. Types: ready, say, tip, ask, answer, summary, context, mood, gesture, fidget, grab, drop, walk, speaking, mute, visible, bubble, command, tour, react, mode, stream, stream-error. Use ask/answer to feed the owner statistics page without storing question text.
Methods
VLSAvatar.init(opts) · say(text, o) · ask(q) · summarize(target, o) · tip(context?)
tour(id) · endTour() · registerTours(map) · point(selector, text, o) · react(event, extra) · setMode(m)
setContext(c) · setMood(m) · gesture(g) · fidget(g?) · lookAt(x,y) · moveTo(x,y) · walkTo(selector|x,y) · walkHome() · resetPosition() · mute(bool) · stop() · show() · hide() · open() · dismiss()
setName(n) · setPersonality({humor, idleChatter, chimeIn, chimeGapMs, rate, pitch, voiceName}) · forget()
registerTips(map) · registerKnowledge(items) · setAskHandler(fn) · setAskEndpoint(url)
setNarrationEndpoint(url, voice) · connectStream(url) · command(cmd) · on(type, fn) · off(type, fn)
getState() · voices() · MOODS · GESTURES · FIDGETS · MODES · REACTIONS