ProjectsA.L.E.K.S.Y v2

A.L.E.K.S.Y v2

A Polish voice assistant that answers in the cloned voice of a real friend.

A Polish-speaking voice assistant for the KNI science club stand at the Maritime University of Szczecin. Say "Aleksy", ask a question, and it answers out loud in the cloned voice of the real Aleksy. A Raspberry Pi 5 handles the wake word, audio and an OLED face, and a Mac mini runs the speech and language models over the network.

Backstory

The first version ran fully offline on an NVIDIA Jetson Xavier NX. It took 20 to 30 seconds to answer, the local model could barely hold a conversation, and the homemade amplifier picked up noise from the board. For v2 I split the work in two: the Pi only does what has to happen in the room, and everything heavy runs on a Mac mini M2 over a WebSocket. I had about a week before the university's adaptation days, so most of it was a race: an OLED that showed a single column until I swapped the power supply, an Audio HAT that took the whole GPIO header, and a 3D-printed case that went through several iterations before it fit. At the event the AI was fine. What failed was the room: the speakers were too quiet for a hall full of stands, and there was no internet. After I got home I added a Bluetooth speaker option and a fallback Wi-Fi hotspot, so the next venue cannot break it the same way.

Problems along the way
  • The OLED showed a single column. It was not the display or the wiring but the power supply. With the official 27 W supply it worked on the first try.
  • The Audio HAT takes the whole GPIO header, so the OLED had to share the I2C bus with the audio codec, and the standoffs between the boards had to go.
  • The case took several iterations. The first did not fit the display or the speakers, the first full print was on bad filament with painful supports, and the M3 holes were too small.
  • The Mac mini M2 is slow next to my M5 Pro laptop. To keep answers under a few seconds, TTS runs at 16 diffusion steps instead of 32 and answers are capped at 300 characters.
  • Speaking sentence by sentence made it worse. It started sooner, but synthesis barely kept up with playback, so there were pauses. The whole answer is now synthesized at once.
  • At the event the speakers were too quiet for a hall full of stands, and with no internet it could not reach the server at all.
Takeaways
  • Splitting the device from the compute was the right call. The Pi stays cheap, small and cool, and the models can change without touching the hardware.
  • Latency matters more than answer quality. A short filler word right after you stop talking does more for how it feels than a better model.
  • A demo device has to work in the worst room, not on the desk at home. Volume and connectivity failed at the event, not the AI.
  • Power problems look like software problems. Rule out the power supply first.
  • Measure on the target hardware early, and give 3D-printed holes some slack: 3.5 mm for M3 screws, not 3.2 mm.
Highlights

Cloned Voice

Answers in the voice of the real Aleksy (OmniVoice)

Wake Word & Filler Words

Says "chwileczkę" right away to cover the wait

OLED Face

Six states: idle, listening, thinking, talking, asleep, error

Own Wi-Fi Hotspot

Pick a network from the panel when no known one is around

Bluetooth Speaker

Selected from the built-in web panel

Screenshots