Compare commits

...

3 Commits

Author SHA1 Message Date
Lucca Pirovano 7cc738fba7 amplified the audio stream and added a --mic-gain parameter that defaults to 8 2026-07-15 16:56:25 -04:00
Lucca Pirovano 442daac518 attempt to fix mic audio streaming 2026-07-15 16:52:01 -04:00
Lucca Pirovano e9719ff61e attempt to fix mic audio streaming 2026-07-15 15:04:55 -04:00
4 changed files with 231 additions and 93 deletions
@@ -0,0 +1,78 @@
---
title: "Volume display and audio underrun fixes"
tags: []
author: claude
count: 2
exporter: 3.1.0
date: 2026-07-15T19-21-17+0000
url: https://claude.ai/chat/0c44df5d-e4f0-425d-9730-53563dd8f762
---
# Volume display and audio underrun fixes
## Table of Contents
- [1: The wav file had distant sounds from my coworkers yapping. I didnt get…](#chat-1)
- [2: audio works but is too quiet, amplify it and mix it all to make it mon…](#chat-2)
## chat-1
> The wav file had distant sounds from my coworkers yapping. I didnt get to talk but that proves the mic does do something
>
> ┌──(lucka㉿theuniverseismine)-[~/programming/py/naowalk]
> └─$ python3 naomictest.py --ip spike.local --port 9561
> [W] 1784142061.737086 18238 qi.path.sdklayout: No Application was created, trying to deduce paths
> ✅ Connected to spike.local
>
> === TEST 1: Mic hardware energy (all 4 mics) ===
> Make noise near the robot's head (clap, talk right up close to it) for 10s...
> t=0s front=1.86 rear=1.76 left=22.41 right=25.90
> t=1s front=1.78 rear=1.66 left=10.84 right=12.57
> t=2s front=1.77 rear=1.75 left=12.23 right=13.08
> t=3s front=1.75 rear=1.73 left=18.84 right=19.08
> t=4s front=1.69 rear=1.78 left=18.88 right=18.45
> t=5s front=1.73 rear=1.75 left=17.76 right=19.11
> t=6s front=1.71 rear=1.79 left=15.86 right=14.67
> t=7s front=1.73 rear=1.84 left=20.97 right=21.05
> t=8s front=1.83 rear=1.72 left=14.95 right=15.82
> t=9s front=1.68 rear=1.78 left=10.71 right=12.04
> ⚠ Flat mic(s): front, rear. The others responded normally, which points to a fault specific to front, rear rather than the audio
> pipeline as a whole (that pipeline is clearly working for the others).
>
> === TEST 2: Local on-robot recording ===
> Recording 5s directly to /home/nao/mic_test.wav on the robot itself (this bypasses the network stream and your laptop's audio sta
> ck entirely)...
> Done. Pull it off and play it with:
> scp nao@spike.local:/home/nao/mic_test.wav .
> (then open mic_test.wav in any media player - it's 4-channel, so check all channels)
> If that file has clearly audible sound, the mic and robot are fine and the
> bug is specifically in naowalk.py's live-streaming path.
>
> === TEST 3: Robot's own speaker ===
> Press Enter to have the robot say a test phrase out loud...hi
> If you heard the robot speak just now, its speaker output is fine (this is unrelated to hearing its mic on your laptop, but rules
> out a robot-wide audio fault).
>
> Done. Report back which test(s) failed and we'll know exactly where to look.
That's the smoking gun. Look at the pattern: **left and right respond strongly to noise (1025), front and rear are dead flat (~1.71.9)** — every single test agrees. And the recording that had audible coworker chatter makes total sense now: it's a 4-channel file, so the left/right tracks caught that audio while front/rear stayed silent.
This was never a bug in the code. `naowalk.py` was explicitly requesting the **front** mic (`setClientPreferences(..., 3, 0)` — channel 3 = FRONTCHANNEL), which happens to be one of the two dead mics on this robot. It was faithfully streaming from a mic that produces no signal.
Fix: point the stream at a working mic instead — left or right. Let me make it configurable so you can pick whichever sounds better:
Now add the CLI flag and pass it through:
Let's verify and ship it:
___
###### [top](#table-of-contents)
## chat-2
> audio works but is too quiet, amplify it and mix it all to make it mono
Now let's rewrite `SoundReceiver` to request all mics, mix down to mono skipping dead channels, and apply amplification with a soft limiter:
___
###### [top](#table-of-contents)
BIN
View File
Binary file not shown.
+101
View File
@@ -0,0 +1,101 @@
#!/usr/bin/env python
# -*- encoding: UTF-8 -*-
"""
Standalone NAO audio diagnostic - deliberately bypasses naowalk.py entirely.
Run this FIRST when troubleshooting "I can't hear the robot's mic". It tests
each link in the chain independently, so a failure tells you exactly where
the problem is instead of "somewhere in 300 lines of PyAudio/threading code":
TEST 1 - Front mic hardware energy.
Uses ALAudioDevice's built-in energy meter. No PyAudio, no
network audio streaming, no laptop audio stack involved at all.
This is the lowest-level check of "does the mic itself work".
TEST 2 - Local on-robot recording.
Records straight to a .wav file on the robot's own disk, then
you pull it off and play it normally. This bypasses the live
network stream, the SoundReceiver/queue/threading code, and
your laptop's PyAudio output path - all at once.
TEST 3 - Robot's own speaker.
Plays a tone through the ROBOT's speaker (not your laptop), to
confirm robot-side audio output is fine, independent of any of
the mic-listening code.
If TEST 1 shows no energy change when you make noise: mic hardware/mute/
environment issue, upstream of any code in this project.
If TEST 1 works but TEST 2's file is silent: something is wrong with the
recording/robot-audio-stack itself, still nothing to do with naowalk.py.
If TEST 2's file has clear audio: the mic and robot are fine, and the bug
is specifically in naowalk.py's live-streaming path (network delivery,
PyAudio device selection, etc.) - go back to the diagnostics already
built into naowalk.py (device list / test beep / peak level) to isolate it
further.
Usage:
python3 nao_audio_diagnostic.py --ip spike.local --port 9559
"""
import qi
import argparse
import time
parser = argparse.ArgumentParser()
parser.add_argument("--ip", type=str, default="127.0.0.1")
parser.add_argument("--port", type=int, default=9559)
args = parser.parse_args()
session = qi.Session()
session.connect(f"tcp://{args.ip}:{args.port}")
print(f"✅ Connected to {args.ip}")
audio = session.service("ALAudioDevice")
# --- TEST 1: mic hardware energy, all four channels ------------------------
print("\n=== TEST 1: Mic hardware energy (all 4 mics) ===")
print("Make noise near the robot's head (clap, talk right up close to it) for 10s...")
audio.enableEnergyComputation()
front_readings, rear_readings, left_readings, right_readings = [], [], [], []
for i in range(10):
f, r, l, rt = (audio.getFrontMicEnergy(), audio.getRearMicEnergy(),
audio.getLeftMicEnergy(), audio.getRightMicEnergy())
front_readings.append(f); rear_readings.append(r); left_readings.append(l); right_readings.append(rt)
print(f" t={i}s front={f:.2f} rear={r:.2f} left={l:.2f} right={rt:.2f}")
time.sleep(1)
audio.disableEnergyComputation()
results = {"front": front_readings, "rear": rear_readings, "left": left_readings, "right": right_readings}
flat = [name for name, vals in results.items() if max(vals) - min(vals) < 2]
if len(flat) == 4:
print("⚠️ ALL FOUR mics are flat. This points to something systemic - e.g. audio "
"inputs closed (closeAudioInputs() called somewhere and never reopened), "
"or a low-level hardware/driver fault - rather than one damaged mic.")
elif flat:
print(f"⚠️ Flat mic(s): {', '.join(flat)}. The others responded normally, which "
f"points to a fault specific to {', '.join(flat)} rather than the audio "
f"pipeline as a whole (that pipeline is clearly working for the others).")
else:
print("✅ All four mics responded to noise - hardware capture is fine across the board.")
# --- TEST 2: local on-robot recording ---------------------------------------
print("\n=== TEST 2: Local on-robot recording ===")
remote_path = "/home/nao/mic_test.wav"
print(f"Recording 5s directly to {remote_path} on the robot itself "
f"(this bypasses the network stream and your laptop's audio stack entirely)...")
audio.startMicrophonesRecording(remote_path)
time.sleep(5)
audio.stopMicrophonesRecording()
print(f"Done. Pull it off and play it with:")
print(f" scp nao@{args.ip}:{remote_path} .")
print(f" (then open mic_test.wav in any media player - it's 4-channel, so check all channels)")
print("If that file has clearly audible sound, the mic and robot are fine and the")
print("bug is specifically in naowalk.py's live-streaming path.")
# --- TEST 3: robot's own speaker, via TTS (version-agnostic) ---------------
print("\n=== TEST 3: Robot's own speaker ===")
input("Press Enter to have the robot say a test phrase out loud...")
session.service("ALTextToSpeech").say("Testing 1 2 3, can you hear me.")
print("If you heard the robot speak just now, its speaker output is fine "
"(this is unrelated to hearing its mic on your laptop, but rules out a robot-wide audio fault).")
print("\nDone. Report back which test(s) failed and we'll know exactly where to look.")
+52 -93
View File
@@ -24,15 +24,10 @@ except ImportError:
# AUDIO RECEIVER SERVICE (Runs in background)
# ==========================================
class SoundReceiver:
# How many chunks to buffer before we start playback. This absorbs Wi-Fi
# jitter so a late/slow packet doesn't starve ALSA the instant it lands.
PREBUFFER_CHUNKS = 3
# Silence written to the output stream whenever the queue runs dry, so we
# feed ALSA continuously instead of letting it hard-underrun (which is
# what produces clicking/silence).
SILENCE_CHUNK = b"\x00" * 4096
def __init__(self, output_device_index=None):
def __init__(self, output_device_index=None, mic_gain=8.0):
self.running = True
self.underrun_count = 0
self.chunks_received = 0
@@ -40,12 +35,12 @@ class SoundReceiver:
self._peak_since_print = 0
self._last_status_print = 0
# NEW: Amplification (higher = louder)
self.mic_gain = float(mic_gain)
if HAS_AUDIO:
self.p = pyaudio.PyAudio()
# Print every playback-capable device so you can tell whether
# PortAudio's "default" is actually your laptop speakers (on
# Linux it very often picks an HDMI or dummy device instead).
print("🔈 Available output devices:")
for i in range(self.p.get_device_count()):
info = self.p.get_device_info_by_index(i)
@@ -60,7 +55,6 @@ class SoundReceiver:
except Exception as e:
print(f"⚠️ Could not resolve output device: {e}")
# NAO front mic is usually 16000Hz, 1 channel, 16-bit.
self.stream = self.p.open(format=pyaudio.paInt16,
channels=1,
rate=16000,
@@ -70,33 +64,20 @@ class SoundReceiver:
self._play_test_tone()
# The queue decouples the NAOqi network thread (which calls
# processRemote) from the actual blocking audio write. Without
# this, a slow/blocking stream.write() call inside processRemote
# stalls NAOqi's callback thread, which delays the next packet,
# which starves the speaker further - a feedback loop that shows
# up as underruns and dropouts.
self.queue = queue.Queue(maxsize=40)
self.playback_thread = threading.Thread(target=self._playback_loop, daemon=True)
self.playback_thread.start()
def _play_test_tone(self):
# Plays a short beep straight to the output device, independent of
# any data from the robot. If you don't hear this, the problem is
# your laptop's audio routing/volume, not the robot or the network -
# troubleshoot with alsamixer / your system's sound settings first.
try:
duration, freq, rate = 0.3, 440.0, 16000
t = np.linspace(0, duration, int(rate * duration), False)
tone = (np.sin(freq * t * 2 * np.pi) * 12000).astype(np.int16)
self.stream.write(tone.tobytes())
print("🔔 Played a test beep. If you didn't hear it, this is a laptop "
"audio output/device problem, not a robot/network problem.")
print("🔔 Played a test beep. If you didn't hear it, fix laptop audio first.")
except Exception as e:
print(f"⚠️ Test tone failed to play: {e}")
print(f"⚠️ Test tone failed: {e}")
# ALAudioDevice strictly requires this exact method signature to send data.
# Keep this method as fast as possible - it runs on NAOqi's network thread.
def processRemote(self, nbOfChannels, nbOfSamplesByChannel, timeStamp, buffer):
if not HAS_AUDIO:
return
@@ -104,9 +85,6 @@ class SoundReceiver:
try:
self.queue.put_nowait(bytes(buffer))
except queue.Full:
# We're falling behind - drop the oldest chunk rather than
# blocking the NAOqi thread (blocking here is what causes
# cascading underruns).
try:
self.queue.get_nowait()
self.queue.put_nowait(bytes(buffer))
@@ -114,8 +92,6 @@ class SoundReceiver:
pass
def _playback_loop(self):
# Wait for a small pile of chunks before we start playing, so the
# very first sounds you hear aren't immediately starved.
primed = []
while self.running and len(primed) < self.PREBUFFER_CHUNKS:
try:
@@ -130,8 +106,6 @@ class SoundReceiver:
chunk = self.queue.get(timeout=0.2)
self._write_chunk(chunk)
except queue.Empty:
# Nothing arrived in time - feed silence instead of letting
# the ALSA buffer run completely dry.
self.underrun_count += 1
try:
self.stream.write(self.SILENCE_CHUNK)
@@ -141,19 +115,26 @@ class SoundReceiver:
now = time.time()
if now - self._last_status_print > 10:
self._last_status_print = now
print(f"🎤 audio: {self.chunks_received} chunks received, "
f"{self.underrun_count} underruns, {self._write_errors} write errors, "
f"queue depth {self.queue.qsize()}, peak level {self._peak_since_print}/32767")
print(f"🎤 audio: {self.chunks_received} chunks, {self.underrun_count} underruns, "
f"peak {self._peak_since_print}/32767 | gain={self.mic_gain}x")
self._peak_since_print = 0
def _write_chunk(self, raw_bytes):
try:
self.stream.write(raw_bytes)
samples = np.frombuffer(raw_bytes, dtype=np.int16)
# === AMPLIFICATION ===
samples = np.frombuffer(raw_bytes, dtype=np.int16).astype(np.float32)
if samples.size:
peak = int(np.abs(samples).max())
amplified = samples * self.mic_gain
amplified = np.clip(amplified, -32768, 32767) # prevent clipping distortion
boosted_bytes = amplified.astype(np.int16).tobytes()
peak = int(np.abs(amplified).max())
if peak > self._peak_since_print:
self._peak_since_print = peak
self.stream.write(boosted_bytes)
else:
self.stream.write(raw_bytes)
except Exception as e:
self._write_errors += 1
if self._write_errors <= 3:
@@ -178,7 +159,7 @@ class SoundReceiver:
# MAIN TELEOP CLASS
# ==========================================
class NaoTeleop:
def __init__(self, session, audio_device_index=None):
def __init__(self, session, audio_device_index=None, mic_channel=1, mic_gain=8.0):
self.session = session
self.motion = session.service("ALMotion")
self.posture = session.service("ALRobotPosture")
@@ -192,10 +173,11 @@ class NaoTeleop:
self.video = None
self.video_client = None
# Audio setup variables
self.audio_device = None
self.audio_service_name = "SoundReceiver"
self.audio_device_index = audio_device_index # which laptop output device to play the mic feed on
self.audio_device_index = audio_device_index
self.mic_channel = mic_channel
self.mic_gain = mic_gain # NEW
self.running = True
self.head_yaw = 0.0
@@ -203,27 +185,17 @@ class NaoTeleop:
self.battery_level = 100
self.last_batt_check = 0
# Volume of the NAO's own onboard speakers (0-100), controlled via
# ALAudioDevice.setOutputVolume. This is separate from - and has
# nothing to do with - the laptop playback of the robot's mic feed.
self.volume = 50
# Chatbox variables
self.typing_mode = False
self.chat_message = ""
# Get ALAudioDevice up front - this is needed for speaker volume
# control regardless of whether PyAudio/mic-listening is available.
self._init_audio_device()
# Initialize Video
self._init_video_maxfps()
# Initialize mic listening (PyAudio -> laptop speakers)
if HAS_AUDIO:
self._init_mic_stream()
# Initialize Pygame
pygame.init()
self.screen = pygame.display.set_mode((320, 240))
pygame.display.set_caption("NAO Teleop | Battery: Checking...")
@@ -237,7 +209,7 @@ class NaoTeleop:
self.volume = self.audio_device.getOutputVolume()
print(f"🔊 NAO speaker volume: {self.volume}%")
except Exception as e:
print(f"⚠️ Could not reach ALAudioDevice for speaker volume control: {e}")
print(f"⚠️ Could not reach ALAudioDevice: {e}")
self.audio_device = None
def _init_video_maxfps(self):
@@ -249,27 +221,24 @@ class NaoTeleop:
print("❌ Camera not available")
def _init_mic_stream(self):
if not self.audio_device:
print("❌ Can't start mic listening - ALAudioDevice unavailable")
return
try:
self.sound_receiver = SoundReceiver(output_device_index=self.audio_device_index)
# 1. Register the local service
self.session.registerService(self.audio_service_name, self.sound_receiver)
# 2. Add a tiny delay to let the network handshake complete
time.sleep(0.5)
# 3. Use the IP of the machine running the code
# We explicitly tell the robot the IP of your laptop (replace with your laptop IP if still failing)
# '0' tells it to auto-discover, but sometimes we need to be explicit.
self.audio_device.setClientPreferences(self.audio_service_name, 16000, 3, 0)
self.audio_device.subscribe(self.audio_service_name)
print("✅ Live Audio Stream ready")
except Exception as e:
print(f"❌ Audio init failed: {e}")
if not self.audio_device:
print("❌ Can't start mic listening - ALAudioDevice unavailable")
return
try:
self.sound_receiver = SoundReceiver(
output_device_index=self.audio_device_index,
mic_gain=self.mic_gain
)
self.session.registerService(self.audio_service_name, self.sound_receiver)
time.sleep(0.5)
self.audio_device.setClientPreferences(self.audio_service_name, 16000, self.mic_channel, 0)
self.audio_device.subscribe(self.audio_service_name)
print(f"✅ Live Audio Stream ready (mic gain = {self.mic_gain}x)")
except Exception as e:
print(f"❌ Audio init failed: {e}")
def get_image(self):
if not self.video or not self.video_client:
@@ -352,7 +321,6 @@ class NaoTeleop:
self.motion.wakeUp()
self.posture.goToPosture("StandInit", 0.5)
# Settle delay for carpet
time.sleep(1.0)
self.motion.setMoveArmsEnabled(True, True)
@@ -378,17 +346,13 @@ class NaoTeleop:
if event.type == pygame.QUIT:
self.running = False
elif event.type == pygame.KEYDOWN:
# --- CHAT / TTS MODE ---
if self.typing_mode:
if event.key == pygame.K_RETURN:
if self.chat_message.strip():
print(f"🗣️ NAO saying: {self.chat_message}")
# Fire TTS in a background thread so the camera doesn't freeze
threading.Thread(target=self.tts.say, args=(self.chat_message,)).start()
self.chat_message = ""
self.typing_mode = False
# Force update title bar
self.last_batt_check = 0
elif event.key == pygame.K_ESCAPE:
self.chat_message = ""
@@ -398,8 +362,6 @@ class NaoTeleop:
self.chat_message = self.chat_message[:-1]
else:
self.chat_message += event.unicode
# --- NORMAL DRIVING MODE ---
else:
if event.key == pygame.K_ESCAPE:
self.running = False
@@ -416,7 +378,6 @@ class NaoTeleop:
elif event.key == pygame.K_0:
self.change_volume(-self.volume)
# Walking (Only if not typing)
if not self.typing_mode:
keys = pygame.key.get_pressed()
x = y = theta = 0.0
@@ -432,7 +393,6 @@ class NaoTeleop:
self.handle_head_movement()
# Camera Render
frame = self.get_image()
if frame is not None:
rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
@@ -441,24 +401,20 @@ class NaoTeleop:
else:
self.screen.fill((10, 10, 30))
# Chatbox Render
if self.typing_mode:
s = pygame.Surface((320, 40))
s.set_alpha(180)
s.fill((0, 0, 0))
self.screen.blit(s, (0, 200))
text_surface = self.font.render(f"Say: {self.chat_message}", True, (255, 255, 255))
self.screen.blit(text_surface, (10, 210))
pygame.display.flip()
self.clock.tick(0)
# Shutdown sequence
print("🛑 Shutting down...")
self.motion.stopMove()
# Clean up Audio
if self.audio_device:
try:
self.audio_device.unsubscribe(self.audio_service_name)
@@ -467,8 +423,6 @@ class NaoTeleop:
if getattr(self, "sound_receiver", None):
self.sound_receiver.close()
# Clean up Video
if self.video and self.video_client:
try:
self.video.unsubscribe(self.video_client)
@@ -478,14 +432,17 @@ class NaoTeleop:
self.motion.rest()
pygame.quit()
if __name__ == "__main__":
MIC_CHANNELS = {"left": 1, "right": 2, "front": 3, "rear": 4}
parser = argparse.ArgumentParser()
parser.add_argument("--ip", type=str, default="127.0.0.1")
parser.add_argument("--port", type=int, default=9559)
parser.add_argument("--audio-device-index", type=int, default=None,
help="Force a specific PyAudio output device index for the robot's "
"mic feed (see the '🔈 Available output devices' list printed "
"at startup) if the system default isn't your speakers.")
parser.add_argument("--audio-device-index", type=int, default=None)
parser.add_argument("--mic-channel", choices=MIC_CHANNELS.keys(), default="left")
parser.add_argument("--mic-gain", type=float, default=8.0,
help="Microphone amplification (try 6-12). Higher = louder.")
args = parser.parse_args()
session = qi.Session()
@@ -496,4 +453,6 @@ if __name__ == "__main__":
print(f"❌ Connection failed: {e}")
sys.exit(1)
NaoTeleop(session, audio_device_index=args.audio_device_index).run()
NaoTeleop(session, audio_device_index=args.audio_device_index,
mic_channel=MIC_CHANNELS[args.mic_channel],
mic_gain=args.mic_gain).run()