Skip to content
OpenCohost Kira

OpenCohost Kira · Active development

OpenCohostKira

An AI co-host powered by PNG/VRM avatars, text-to-speech (TTS), memory, profiles, an agenda that lets the LLM iterate on itself, and Twitch stream chat ingestion.

Kira, the OpenCohost AI co-host, presenting the product
  • Turn-based conversation

    Requests are queued and answered in short turns instead of overlapping.

  • Voice in the runtime capture

    Hear Kira's voice in the current product capture.

  • Failure recovery

    See how the current interface recovers from an interaction failure.

Watch Kira in action

Uncut runtime video: turn-based voice and supervised recovery.

    • Host-controlled segment

      Interrupt, redirect, or stop Kira at any time.

    • Useful chat context

      Spam and spikes become one co-host reaction.

    • Queued conversation

      Short turns prevent overlap.

    • Voice through LiveAudio

      Push-to-talk and transcribed speech feed clean context to Kira.

    • A home for its voice

      The model — local or cloud — gets its own speech output: local synthesis with Piper or light synthesis with Edge-TTS.

    • Guardrails always on

      Unsafe instructions are blocked before they reach the stream.

Audio: Spanish · English capture in progress

Preview upcoming v0.5.0-alpha
Aug – Oct 2026 development

Preview the upcoming Kira v0.5.0-alpha

An advance look at the next major release: dedicated 3D VRM expressions, desktop companion overlay, integrated model management, and our local MiniLM v5 memory pipeline.

Star or follow the repository to track upcoming releasesStar on GitHub

Preview capture: desktop companion, lip-sync visemes, and in-app model downloader

Key v0.5.0 highlights

01

Dedicated VRM & Expression Engine

Custom 3D model designed specifically for Kira, featuring synchronized lip-sync and audio visemes.

02

Desktop Companion Mode

Interact with Kira in an independent, floating companion overlay synchronized with the main timeline.

03

In-App Model Downloader

Search and pull Ollama models directly from the application interface.

04

VRAM-Aware Context & Adaptive Reasoning

Automatic context scaling up to 16k tokens, spill protection, and adaptive reasoning budgets with manual ceilings.

05

Flexible Profile System

Clean, modular persona prompts removing legacy hardcoded constraints.

06

Memory v5

End-to-end all-MiniLM embedding integration for improved semantic indexing and recall. Local embedding — MiniLM-L12 ONNX served by an isolated background worker (no cloud, no extra VRAM on the LLM runner).

07

Speech Flow Controls

Option to discard audio residues on interruption and purge stale speech on fast regenerations.

08

/setup command

First-run LLM configuration (provider, Ollama, models) from the chat palette.

OpenCohost Kira desktop interface showing a blonde VRM avatar named AlKira, with an active chat panel and a live conversation. Status bar shows local gemma model running.
OpenCohost Kira desktop interface with a blonde avatar, expanded sidebar with profile list, and a Spanish conversation about Hades 2. Status: local co-host, inactive.
OpenCohost Kira: avatar and supervised conversation. Interface shown in Spanish.

A supervised co-host for your live show.

Kira brings your avatar, chat context and prepared agenda together; you stay in control.

Shape Kira with editable profiles.

Edit profile dialog in OpenCohost Kira showing name, system prompt, extended responses toggle, recent memory turns set to 10, and delete profile. Interface shown in Spanish.
Edit a profile’s prompt and memory settings. Interface shown in Spanish.

4 h 27 min of supervised agenda flow.

One instrumented run; not proof of production streaming or autonomous hosting.

Exact duration 4:27:35
4 h 27 min
agenda session
Traced generations
285
runtime total
Cloud / local generations
120 / 165
provider split

Is your setup ready for Kira?

Local performance depends on your GPU, free VRAM, model, and game load.

RTX 3060 12GB

Game Type
Light / Competitive (Fortnite, Valorant, LoL)
Experience
✅ Light voice mode; responsive co-hosting.Tested

RTX 3060 12GB

Game Type
Heavy / Ultra 3D (Cyberpunk, Alan Wake 2)
Experience
⚠️ Light voice mode recommended; minor delay possible.Estimate

RTX 4090

Game Type
Any game / settings
Experience
✅ More headroom for heavy modes; depends on model, game load, and features.Estimate

GTX 1660

Game Type
Lightweight games (Fortnite, Valorant)
Experience
✅ Light voice; disable high-poly avatar.Estimate

No Dedicated GPU

Game Type
Just chatting screen / browsing
Experience
❌ Local voice processing needs a dedicated GPU.Estimate

Only one PC is tested: RTX 3060 12GB. All other rows are estimates, not benchmarks; results depend on model, GPU, free VRAM, and game load.

Download the public alpha for Windows.

Windows alpha, published from CI. No hosted service or bundled model: bring Ollama, your chosen model, and suitable hardware.

v0.3.0-alpha.4 · 588 MB · Windows 10/11 x64 · published Sep 9, 2026 · Verify the SHA256 checksums

CI-built alpha. Report issues without API keys, private logs, or sensitive information.

Help improve the alpha.

Share reproducible setup or workflow issues. Never include API keys, private logs, or sensitive information.

Share feedback

Three honest answers before you evaluate Kira

What is Kira today?

An alpha streaming co-host with supervised agenda, Direct/PTT, chat, memory and local/cloud model paths. It is not an autonomous host.

Can I use local or cloud models?

Both paths exist. Ollama has been tested with specific Gemma, Llama and Qwen configurations; BYOK supports cloud providers. Results depend on the exact model and hardware.

Is streaming production-ready?

No. It has not been personally validated in a real stream — only with external streamers by consuming their chats, and only as chat consumption, not as a proven pleasant experience for viewers. Twitch compatibility follows Twitch policies. YouTube is a different story: reading its chat uses a method YouTube Terms of Service do not allow, so using Kira there carries a real risk of a channel ban. In the app, live viewer chat is read from the Stream tab and remains in development: LLM responses can be inaccurate or inappropriate depending on your active audience profile, and whatever Kira says on your channel is your responsibility. If your LLM provider is cloud-based instead of local, viewer chat content leaves your machine.