Compare commits
25 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| ae9d1652a4 | |||
| 7d0b968b36 | |||
| 4a4625028b | |||
| 5977e6fc7d | |||
| 6a97da3b86 | |||
| 17c7da2125 | |||
| b30fc10648 | |||
| f806fa8efc | |||
| e1b5e3ab76 | |||
| 929687dea5 | |||
| a6767bf154 | |||
| f4140328e6 | |||
| c6b9cb0718 | |||
| ca196f5f1e | |||
| 158372e528 | |||
| ab67b128c1 | |||
| cdb2e1f90f | |||
| 99f2d6a9cd | |||
| bf3e197a20 | |||
| a5cdb4e4bd | |||
| c69f8ca7d0 | |||
| 043c4c3f09 | |||
| 58fc2af593 | |||
| 34305fbeea | |||
| e943073717 |
@@ -0,0 +1,4 @@
|
||||
#!/bin/bash
|
||||
pactl set-source-volume alsa_input.pci-0000_00_1b.0.analog-stereo 32%
|
||||
#python3.11 naowalk.py --ip spike.local --port 9561 --mic-gain 9 --video-scale 3.0
|
||||
python3.11 naowalk.py --ip 127.0.0.1 --port 9561 --mic-gain 9 --video-scale 3.0 --nao-ssh-port 2222 --media-relay 127.0.0.1:8000
|
||||
@@ -0,0 +1,105 @@
|
||||
# naowalk
|
||||
|
||||
A keyboard teleop client for the NAO robot: live video, walking, head control,
|
||||
two-way audio (hear the robot's mics, make it speak via TTS), and a YouTube
|
||||
music player that downloads a track with `yt-dlp` and plays it through the
|
||||
robot's own speakers.
|
||||
|
||||
## Requirements
|
||||
|
||||
**System packages** (Debian/Ubuntu names — adjust for your OS):
|
||||
|
||||
```bash
|
||||
sudo apt install portaudio19-dev ffmpeg sshpass openssh-client
|
||||
```
|
||||
|
||||
- `portaudio19-dev` — needed to build/install `pyaudio`
|
||||
- `ffmpeg` — needed by `yt-dlp` to extract audio to mp3
|
||||
- `sshpass` + `openssh-client` — used to `scp` downloaded tracks onto the
|
||||
robot and to run `mkdir` over `ssh` (the script assumes the default
|
||||
`nao`/`nao` credentials)
|
||||
|
||||
**NAOqi Python SDK (`qi`)** — not on PyPI. Download the `pynaoqi-python`
|
||||
wheel that matches your OS/Python version from the Aldebaran/SoftBank
|
||||
Robotics developer site and install it:
|
||||
|
||||
```bash
|
||||
pip install ./pynaoqi-*.whl
|
||||
```
|
||||
|
||||
**Python packages:**
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
python3 naowalk.py --ip <robot-ip> [options]
|
||||
```
|
||||
|
||||
| Flag | Default | Description |
|
||||
|---|---|---|
|
||||
| `--ip` | `127.0.0.1` | Robot's IP address |
|
||||
| `--port` | `9559` | NAOqi port |
|
||||
| `--audio-device-index` | auto | Output device index for hearing the robot's mic on your machine |
|
||||
| `--mic-channel` | `left` | Which robot mic to stream (`left`, `right`, `front`, `rear`) — pick a working one, see Troubleshooting |
|
||||
| `--mic-gain` | `8.0` | Amplification applied to the streamed mic audio |
|
||||
| `--video-scale` | `2.0` | Display window scale (base feed is 320x240) |
|
||||
|
||||
## Controls
|
||||
|
||||
| Key | Action |
|
||||
|---|---|
|
||||
| `W`/`A`/`S`/`D` or arrows | Walk forward/turn/back |
|
||||
| `I`/`K`/`J`/`L` | Head pitch/yaw |
|
||||
| `Space` | Reset head to center |
|
||||
| `1` | Wave gesture |
|
||||
| `M` | Open music search (type query, `Enter` to search, `↑↓` to select, `Enter`/`1`-`5` to play) |
|
||||
| `P` | Pause / resume music (resume restarts the track from the beginning — see Notes) |
|
||||
| `X` | Stop music |
|
||||
| `-` / `=` | Volume down/up |
|
||||
| `T` | Type a message for the robot to speak (TTS) |
|
||||
| `Esc` | Quit (or cancel current mode) |
|
||||
|
||||
## How it works
|
||||
|
||||
- **`SoundReceiver`** — registered as a NAOqi service that receives the
|
||||
robot's mic stream and plays it out your local speakers via `pyaudio`, with
|
||||
gain and a prebuffer to smooth out network jitter.
|
||||
- **`NaoMusicPlayer`** — searches YouTube with `yt-dlp`, downloads and
|
||||
converts the top match to mp3, `scp`s it to the robot, and plays it with
|
||||
`ALAudioPlayer`.
|
||||
- **`NaoTeleop`** — the main loop: pygame window for video/input, keyboard
|
||||
handling, head/walk control, and wiring everything above together.
|
||||
|
||||
## Notes / design decisions
|
||||
|
||||
- **No seek.** Seeking was removed entirely — it added the most complexity
|
||||
for the least-used feature. `P` is a simple play/stop toggle; resuming
|
||||
restarts the track from the beginning rather than tracking a playback
|
||||
position.
|
||||
- **`moveInit()` on startup.** The robot used to need its walk buttons
|
||||
jittered for a few seconds before its first walk would work. `run()` now
|
||||
calls `ALMotionProxy::moveInit()` once at startup, which settles the walk
|
||||
engine into the right pose before any real move command — this is the
|
||||
documented NAOqi fix for exactly this symptom.
|
||||
- **`_safe()` helper.** Most one-off NAOqi calls (setting angles, volume,
|
||||
stopping movement, unsubscribing services) are wrapped in a small
|
||||
`self._safe(label, fn)` helper instead of a repeated `try/except` block, so
|
||||
one failed call logs and moves on without crashing the teleop loop or the
|
||||
shutdown sequence.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **Mic sounds dead / silent on one channel:** some robots have one or more
|
||||
physical mics that produce no signal. Test all four with a quick energy
|
||||
check, then pick a working channel with `--mic-channel`.
|
||||
- **Robot falls over on the very first walk command:** make sure you're on a
|
||||
build that includes the `moveInit()` startup call described above; if it's
|
||||
still happening, increase the `time.sleep(1.0)` after `goToPosture` to give
|
||||
the posture more time to settle first.
|
||||
- **`scp`/`ssh` failures for music:** confirm the robot is reachable at
|
||||
`--ip` and that `nao`/`nao` are still the correct credentials; the script
|
||||
hardcodes them via `sshpass`.
|
||||
Binary file not shown.
+41
@@ -0,0 +1,41 @@
|
||||
---
|
||||
title: "Refactoring Python project files and code"
|
||||
tags: []
|
||||
author: claude
|
||||
count: 2
|
||||
exporter: 3.1.0
|
||||
date: 2026-07-16T17-38-58+0000
|
||||
url: https://claude.ai/chat/313cfd88-e6b9-4256-b155-0b389d2b86c4
|
||||
---
|
||||
|
||||
# Refactoring Python project files and code
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [1: make a requirements.txt, readme.md, and also remove the seek function…](#chat-1)
|
||||
- [2: how much faster can i make it walk, can we make it adjustable during r…](#chat-2)
|
||||
|
||||
## chat-1
|
||||
|
||||
> make a requirements.txt, readme.md, and also remove the seek function completely and reduce the lines of code without hurting stability
|
||||
|
||||
Now the `pause()` method — simplify to a plain stop/restart toggle (no seek, no position math):
|
||||
|
||||
Now let's add the `_safe` helper to `NaoTeleop` and use it to collapse the repeated try/except boilerplate:
|
||||
|
||||
Now the shutdown block — this is where `_safe` pays off the most:
|
||||
|
||||
Good — that's stable. Now let's view the final file once more end-to-end to make sure everything reads cleanly before shipping it.
|
||||
|
||||
This looks clean and correct. Now let's write the `requirements.txt` and `README.md`.
|
||||
|
||||
Good, all confirmed. Now the README:
|
||||
|
||||
Now let's copy everything to outputs and present the files:
|
||||
|
||||
___
|
||||
###### [top](#table-of-contents)
|
||||
## chat-2
|
||||
|
||||
> how much faster can i make it walk, can we make it adjustable during runtime? we might need to change its posture so its less likely to fall and more balanced for carpet
|
||||
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
---
|
||||
title: "Replacing music player backend with yt-dlp"
|
||||
tags: []
|
||||
author: claude
|
||||
count: 2
|
||||
exporter: 3.1.0
|
||||
date: 2026-07-16T17-20-34+0000
|
||||
url: https://claude.ai/chat/8bb5e420-7ad2-4048-be3f-95894713aff6
|
||||
---
|
||||
|
||||
# Replacing music player backend with yt-dlp
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [1: lets just gut the seek function and focus on what works. we also need…](#chat-1)
|
||||
- [2: alright awesome, lets remove the start from current position also, i w…](#chat-2)
|
||||
|
||||
## chat-1
|
||||
|
||||
> lets just gut the seek function and focus on what works. we also need to address the robot falling over if i try to walk from standstill, but only the first time i do it. If i jitter the walk forward and back buttons for about 3 seconds to get over the initial walk issues, it can start walking without falling over. how do we fix this
|
||||
|
||||
Both very solvable. Let's do the walk fix first since it's the more important one — confirmed via NAOqi docs: `ALMotionProxy::moveInit()` is the standard fix for exactly this. It's a one-time blocking call that initializes the walk engine and settles the robot into the right pose _before_ the first real move command, which is exactly what your button-jittering was accidentally doing by brute force. It's missing from the script entirely.
|
||||
|
||||
Now let's gut `seek()` entirely per your request — but keep the position tracking it shares with `pause()`, since that's still needed for resume-from-position to work:
|
||||
|
||||
Right, orphaned body left over. Let's clean that up:
|
||||
|
||||
Now remove the comma/period key bindings that called `seek()`:
|
||||
|
||||
Now compile-check and confirm nothing else references `seek`:
|
||||
|
||||
Clean. Let's ship it:
|
||||
|
||||
___
|
||||
###### [top](#table-of-contents)
|
||||
## chat-2
|
||||
|
||||
> alright awesome, lets remove the start from current position also, i want to get the lines of code down, what can be done without compromising on the stability and functionality to simplify our code
|
||||
|
||||
Now let's do the full pass. First, the concrete simplification you asked for — remove position tracking, revert pause/resume to restart-from-beginning:
|
||||
|
||||
Now `pause()` becomes a plain stop/restart toggle — no position math, no `playFileFromPosition`:
|
||||
|
||||
Now let's cut the repetitive `try/except Exception as e: print(...)` boilerplate scattered through `NaoTeleop` — same safety, less code, via one small helper:
|
||||
|
||||
Now let's add the `_safe` helper and apply it across these repeated sites:
|
||||
|
||||
Now the `moveInit` startup call:
|
||||
|
||||
Now the shutdown block — this is where it pays off most:
|
||||
|
||||
Let's check the movement block and the `T` key's `stopMove()` call too:
|
||||
|
||||
___
|
||||
###### [top](#table-of-contents)
|
||||
|
||||
@@ -0,0 +1,53 @@
|
||||
---
|
||||
title: "Your first chat with Claude"
|
||||
tags: []
|
||||
author: claude
|
||||
count: 5
|
||||
exporter: 3.1.0
|
||||
date: 2026-07-16T20-40-30+0000
|
||||
url: https://claude.ai/chat/838b001f-8bf8-4539-98ec-467513fb9860
|
||||
---
|
||||
|
||||
# Your first chat with Claude
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [1: add the 6 7 meme as a third gesture](#chat-1)
|
||||
- [2: kinda looks like hes flailing his arms around, isnt it just the elbows…](#chat-2)
|
||||
- [3: right motions, but the nao currently extends his arms out, then moves…](#chat-3)
|
||||
- [4: perfect, now just rotate the hands so the palm faces up, everything el…](#chat-4)
|
||||
- [5: make him tts say six seven, also remove the input delay on starting th…](#chat-5)
|
||||
|
||||
## chat-1
|
||||
|
||||
> add the 6 7 meme as a third gesture
|
||||
|
||||
## chat-2
|
||||
|
||||
> kinda looks like hes flailing his arms around, isnt it just the elbows moving in the meme? with the arms in front
|
||||
|
||||
Now let me replace the whole method with a version where the shoulders lock the arms up in front and only the elbows do the seesaw motion.
|
||||
|
||||
___
|
||||
###### [top](#table-of-contents)
|
||||
## chat-3
|
||||
|
||||
> right motions, but the nao currently extends his arms out, then moves his elbows to push the forearms close to his chest. i think the rotation is wrong on the elbow
|
||||
|
||||
That makes sense — at `ElbowYaw=0` the elbow's bend axis is oriented so flexing curls the forearm in across the body toward the chest, not up/down in place. Rotating `ElbowYaw` about 90° turns that same hinge into a front-facing "tip" instead, and I'll shrink the swing so it's a wobble, not a full extend-then-fold.
|
||||
|
||||
___
|
||||
###### [top](#table-of-contents)
|
||||
## chat-4
|
||||
|
||||
> perfect, now just rotate the hands so the palm faces up, everything else is perfect
|
||||
|
||||
## chat-5
|
||||
|
||||
> make him tts say six seven, also remove the input delay on starting the 3rd gesture
|
||||
|
||||
`self.tts.say(...)` is a blocking call, so it was speaking first and only starting the arm motion once the phrase finished — that's the delay. Making it fire on its own thread lets the gesture start immediately while the speech plays alongside it.
|
||||
|
||||
___
|
||||
###### [top](#table-of-contents)
|
||||
|
||||
@@ -0,0 +1,327 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
"""
|
||||
nao_video_server.py - RUNS ON THE ROBOT, under its own Python 2.7.
|
||||
|
||||
Opens a qi session to NAOqi over localhost (zero network hop - it's
|
||||
the same box) and serves TWO persistent, length-framed TCP streams -
|
||||
one for compressed video (JPEG), one for compressed audio (mu-law) -
|
||||
each on its own port/connection. They used to share one connection,
|
||||
but a single big video frame mid-send could block audio behind it for
|
||||
its whole transfer time, which is a bad tradeoff for something as
|
||||
latency-sensitive as audio. Splitting them means a slow video frame
|
||||
can never delay audio.
|
||||
|
||||
Why not HTTP: the phone tunnel only forwards specific TCP ports, so
|
||||
this still rides over TCP - but a 5-byte binary header has none of
|
||||
HTTP's chunked/multipart-boundary parsing overhead, and it's trivial
|
||||
to keep alive and reconnect on a flaky link.
|
||||
|
||||
Why mic capture lives here now: it used to be shipped raw (16-bit PCM,
|
||||
uncompressed) straight through NAOqi's own pub/sub over the qi RPC
|
||||
session the client holds - i.e. across the tunnel, uncompressed, the
|
||||
whole time. Subscribing to ALAudioDevice locally (same zero-hop trick
|
||||
as the camera) means only mu-law-compressed bytes ever cross the
|
||||
tunnel, half the size of raw PCM.
|
||||
|
||||
Wire format, identical on both connections (all integers big-endian,
|
||||
via `struct`):
|
||||
1 byte type 0x00 HELLO 0x01 VIDEO 0x02 AUDIO
|
||||
4 bytes length
|
||||
N bytes payload
|
||||
HELLO payload: JSON, informational only
|
||||
VIDEO payload: JPEG bytes (video connection only)
|
||||
AUDIO payload: mu-law encoded 8-bit samples, mono (audio connection only)
|
||||
|
||||
Under load (slow/lossy tunnel), video quality and frame rate step
|
||||
down automatically (see QUALITY_TIERS) and recover once sends are
|
||||
fast again - and since encoding is throttled to match, this also cuts
|
||||
CPU use on a congested link instead of just wasting it on frames that
|
||||
would've been discarded anyway.
|
||||
|
||||
Needs the SDK's own env vars, since the bindings aren't on a login
|
||||
shell's default path:
|
||||
|
||||
export PYTHONPATH=/opt/aldebaran/lib/python2.7/site-packages
|
||||
export LD_LIBRARY_PATH=/opt/aldebaran/lib
|
||||
python2 nao_video_server.py
|
||||
|
||||
Stdlib only (BaseHTTPServer/Flask are gone entirely now) plus qi,
|
||||
cv2, and numpy, which are already confirmed present on the robot.
|
||||
"""
|
||||
import collections
|
||||
import json
|
||||
import socket
|
||||
import struct
|
||||
import threading
|
||||
import time
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import qi
|
||||
|
||||
# ---- tunables ----
|
||||
NAOQI_PORT = 9561 # local qi session port (confirmed working)
|
||||
VIDEO_PORT = 8000 # forward both of these through the phone tunnel
|
||||
AUDIO_PORT = 8001 # kept on its own connection so a slow video
|
||||
# frame can never block/delay audio delivery
|
||||
CAMERA_RESOLUTION = 1 # 0=kQQVGA(160x120) 1=kQVGA(320x240) 2=kVGA(640x480)
|
||||
CAMERA_COLORSPACE = 11 # kRGBColorSpace
|
||||
CAMERA_FPS = 20
|
||||
AUDIO_RATE = 16000
|
||||
AUDIO_CHANNELS = 1
|
||||
MIC_CHANNEL = 1 # 1=left 2=right 3=front 4=rear (ALAudioDevice channel index)
|
||||
AUDIO_QUEUE_MAX = 60 # chunks buffered before we start dropping the oldest
|
||||
|
||||
# (jpeg_quality, min_seconds_between_frames) - best link first, worst last.
|
||||
# The server steps down a tier when a video send is slow, and back up
|
||||
# after a run of fast ones.
|
||||
QUALITY_TIERS = [
|
||||
(70, 0.05),
|
||||
(55, 0.07),
|
||||
(40, 0.10),
|
||||
(25, 0.15),
|
||||
]
|
||||
SLOW_SEND_THRESHOLD = 0.20 # seconds - a send this slow means congestion
|
||||
FAST_STREAK_TO_RECOVER = 30 # consecutive quick sends before stepping back up
|
||||
|
||||
TYPE_HELLO = 0x00
|
||||
TYPE_VIDEO = 0x01
|
||||
TYPE_AUDIO = 0x02
|
||||
|
||||
|
||||
def encode_ulaw(pcm16):
|
||||
"""int16 numpy array -> uint8 numpy array, mu-law companded.
|
||||
Must match decode_ulaw() in naowalk.py exactly, or the far end
|
||||
just hears noise - keep the two in sync if you touch this."""
|
||||
mu = 255.0
|
||||
x = np.clip(pcm16.astype(np.float64) / 32768.0, -1.0, 1.0)
|
||||
y = np.sign(x) * np.log1p(mu * np.abs(x)) / np.log1p(mu)
|
||||
return (((y + 1.0) / 2.0) * 255.0).round().astype(np.uint8)
|
||||
|
||||
|
||||
class FrameGrabber(object):
|
||||
"""Background thread: keeps the single most recent JPEG-encoded
|
||||
frame available, at whatever quality tier the link currently
|
||||
supports. A stalled capture never blocks a client's read."""
|
||||
|
||||
def __init__(self, session):
|
||||
self.video = session.service("ALVideoDevice")
|
||||
self.client = self.video.subscribeCamera(
|
||||
"MediaServer", 0, CAMERA_RESOLUTION, CAMERA_COLORSPACE, CAMERA_FPS)
|
||||
self._lock = threading.Lock()
|
||||
self._jpeg = None
|
||||
self._tier = 0
|
||||
self._fast_streak = 0
|
||||
self._stop = threading.Event()
|
||||
self._thread = threading.Thread(target=self._loop)
|
||||
self._thread.daemon = True
|
||||
self._thread.start()
|
||||
|
||||
def _loop(self):
|
||||
last_encode = 0.0
|
||||
while not self._stop.is_set():
|
||||
try:
|
||||
image = self.video.getImageRemote(self.client)
|
||||
if not image or len(image) < 7:
|
||||
continue
|
||||
# getImageRemote() blocks until the next camera frame is
|
||||
# ready (~1/CAMERA_FPS), so this loop is already paced -
|
||||
# but under a congested tier we send far less often than
|
||||
# that, so skip the (expensive) convert+encode work for
|
||||
# frames we're just going to overwrite unsent anyway.
|
||||
quality, min_interval = self.quality_and_interval()
|
||||
now = time.time()
|
||||
if now - last_encode < min_interval:
|
||||
continue
|
||||
w, h = image[0], image[1]
|
||||
arr = np.frombuffer(bytearray(image[6]), dtype=np.uint8).reshape((h, w, 3))
|
||||
bgr = cv2.cvtColor(arr, cv2.COLOR_RGB2BGR)
|
||||
ok, buf = cv2.imencode(".jpg", bgr, [cv2.IMWRITE_JPEG_QUALITY, quality])
|
||||
if ok:
|
||||
with self._lock:
|
||||
self._jpeg = buf.tobytes()
|
||||
last_encode = now
|
||||
except Exception as e:
|
||||
print("frame grab error: %s" % e)
|
||||
time.sleep(0.2)
|
||||
|
||||
def quality_and_interval(self):
|
||||
with self._lock:
|
||||
return QUALITY_TIERS[self._tier]
|
||||
|
||||
def report_send_time(self, elapsed):
|
||||
"""Called by the writer thread after each video send so we can
|
||||
adapt to how the link is actually behaving right now."""
|
||||
with self._lock:
|
||||
if elapsed > SLOW_SEND_THRESHOLD:
|
||||
self._fast_streak = 0
|
||||
if self._tier < len(QUALITY_TIERS) - 1:
|
||||
self._tier += 1
|
||||
print("link looks congested - dropping to quality tier %d" % self._tier)
|
||||
else:
|
||||
self._fast_streak += 1
|
||||
if self._fast_streak >= FAST_STREAK_TO_RECOVER and self._tier > 0:
|
||||
self._tier -= 1
|
||||
self._fast_streak = 0
|
||||
print("link recovered - raising to quality tier %d" % self._tier)
|
||||
|
||||
def latest(self):
|
||||
with self._lock:
|
||||
return self._jpeg
|
||||
|
||||
def close(self):
|
||||
self._stop.set()
|
||||
try:
|
||||
self.video.unsubscribe(self.client)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
class MicGrabber(object):
|
||||
"""Subscribes to the robot's own mic locally (zero-hop, same trick
|
||||
as FrameGrabber) and mu-law-encodes each chunk as it arrives. This
|
||||
replaces the old design where raw PCM was pushed to the client
|
||||
directly through NAOqi's pub/sub over the (tunneled) qi session."""
|
||||
|
||||
def __init__(self, session, channel=MIC_CHANNEL):
|
||||
self.audio = session.service("ALAudioDevice")
|
||||
self._svc_name = "MicRelay"
|
||||
self._lock = threading.Lock()
|
||||
self._chunks = collections.deque(maxlen=AUDIO_QUEUE_MAX)
|
||||
session.registerService(self._svc_name, self)
|
||||
time.sleep(0.5) # let NAOqi's service directory propagate the
|
||||
# registration before ALAudioDevice looks it up
|
||||
# by name, or subscribe() fails to find it
|
||||
self.audio.setClientPreferences(self._svc_name, AUDIO_RATE, channel, 0)
|
||||
self.audio.subscribe(self._svc_name)
|
||||
|
||||
def processRemote(self, nbOfChannels, nbOfSamplesByChannel, timeStamp, buffer):
|
||||
try:
|
||||
pcm = np.frombuffer(bytes(buffer), dtype=np.int16)
|
||||
chunk = encode_ulaw(pcm).tobytes()
|
||||
with self._lock:
|
||||
self._chunks.append(chunk)
|
||||
except Exception as e:
|
||||
print("mic encode error: %s" % e)
|
||||
|
||||
def pop_all(self):
|
||||
"""Drain everything queued right now, oldest first."""
|
||||
with self._lock:
|
||||
out = list(self._chunks)
|
||||
self._chunks.clear()
|
||||
return out
|
||||
|
||||
def close(self):
|
||||
try:
|
||||
self.audio.unsubscribe(self._svc_name)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def _send(conn, msg_type, payload):
|
||||
conn.sendall(struct.pack(">BI", msg_type, len(payload)))
|
||||
conn.sendall(payload)
|
||||
|
||||
|
||||
def serve_video(grabber, host, port):
|
||||
"""Video gets its own connection/thread - a slow sendall() here
|
||||
(a big JPEG frame over a bad link) must never be able to delay
|
||||
audio, which used to share this same socket and paid for it in
|
||||
latency."""
|
||||
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
||||
sock.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
|
||||
sock.bind((host, port))
|
||||
sock.listen(1)
|
||||
print("video server listening on :%d" % port)
|
||||
while True:
|
||||
conn, addr = sock.accept()
|
||||
conn.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
|
||||
conn.settimeout(5.0)
|
||||
print("video client connected: %s" % (addr,))
|
||||
last_jpeg = None
|
||||
last_send = 0.0
|
||||
try:
|
||||
_send(conn, TYPE_HELLO, json.dumps({"stream": "video"}).encode("utf-8"))
|
||||
while True:
|
||||
jpeg = grabber.latest()
|
||||
_, min_interval = grabber.quality_and_interval()
|
||||
now = time.time()
|
||||
if jpeg is not None and jpeg is not last_jpeg and now - last_send >= min_interval:
|
||||
t0 = time.time()
|
||||
_send(conn, TYPE_VIDEO, jpeg)
|
||||
grabber.report_send_time(time.time() - t0)
|
||||
last_jpeg = jpeg
|
||||
last_send = now
|
||||
else:
|
||||
time.sleep(0.01)
|
||||
except Exception as e:
|
||||
print("video client disconnected (%s)" % e)
|
||||
finally:
|
||||
try:
|
||||
conn.close()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def serve_audio(mic, host, port):
|
||||
"""Audio's own connection/thread. Checked eagerly (short sleep,
|
||||
not paced like video) since gaps are far more noticeable than a
|
||||
video frame being a beat stale."""
|
||||
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
|
||||
sock.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
|
||||
sock.bind((host, port))
|
||||
sock.listen(1)
|
||||
print("audio server listening on :%d" % port)
|
||||
while True:
|
||||
conn, addr = sock.accept()
|
||||
conn.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
|
||||
conn.settimeout(5.0)
|
||||
print("audio client connected: %s" % (addr,))
|
||||
try:
|
||||
_send(conn, TYPE_HELLO,
|
||||
json.dumps({"rate": AUDIO_RATE, "channels": AUDIO_CHANNELS}).encode("utf-8"))
|
||||
while True:
|
||||
chunks = mic.pop_all() if mic is not None else []
|
||||
if chunks:
|
||||
for chunk in chunks:
|
||||
_send(conn, TYPE_AUDIO, chunk)
|
||||
else:
|
||||
time.sleep(0.005)
|
||||
except Exception as e:
|
||||
print("audio client disconnected (%s)" % e)
|
||||
finally:
|
||||
try:
|
||||
conn.close()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def main():
|
||||
session = qi.Session()
|
||||
session.connect("tcp://127.0.0.1:%d" % NAOQI_PORT)
|
||||
print("connected to local NAOqi on port %d" % NAOQI_PORT)
|
||||
|
||||
grabber = FrameGrabber(session)
|
||||
mic = None
|
||||
try:
|
||||
mic = MicGrabber(session)
|
||||
print("mic capture ready (mu-law, %dHz)" % AUDIO_RATE)
|
||||
except Exception as e:
|
||||
print("mic capture unavailable (%s) - video only" % e)
|
||||
|
||||
video_thread = threading.Thread(target=serve_video, args=(grabber, "0.0.0.0", VIDEO_PORT))
|
||||
video_thread.daemon = True
|
||||
video_thread.start()
|
||||
|
||||
try:
|
||||
serve_audio(mic, "0.0.0.0", AUDIO_PORT)
|
||||
except KeyboardInterrupt:
|
||||
pass
|
||||
finally:
|
||||
grabber.close()
|
||||
if mic:
|
||||
mic.close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
-101
@@ -1,101 +0,0 @@
|
||||
#!/usr/bin/env python
|
||||
# -*- encoding: UTF-8 -*-
|
||||
"""
|
||||
Standalone NAO audio diagnostic - deliberately bypasses naowalk.py entirely.
|
||||
|
||||
Run this FIRST when troubleshooting "I can't hear the robot's mic". It tests
|
||||
each link in the chain independently, so a failure tells you exactly where
|
||||
the problem is instead of "somewhere in 300 lines of PyAudio/threading code":
|
||||
|
||||
TEST 1 - Front mic hardware energy.
|
||||
Uses ALAudioDevice's built-in energy meter. No PyAudio, no
|
||||
network audio streaming, no laptop audio stack involved at all.
|
||||
This is the lowest-level check of "does the mic itself work".
|
||||
|
||||
TEST 2 - Local on-robot recording.
|
||||
Records straight to a .wav file on the robot's own disk, then
|
||||
you pull it off and play it normally. This bypasses the live
|
||||
network stream, the SoundReceiver/queue/threading code, and
|
||||
your laptop's PyAudio output path - all at once.
|
||||
|
||||
TEST 3 - Robot's own speaker.
|
||||
Plays a tone through the ROBOT's speaker (not your laptop), to
|
||||
confirm robot-side audio output is fine, independent of any of
|
||||
the mic-listening code.
|
||||
|
||||
If TEST 1 shows no energy change when you make noise: mic hardware/mute/
|
||||
environment issue, upstream of any code in this project.
|
||||
If TEST 1 works but TEST 2's file is silent: something is wrong with the
|
||||
recording/robot-audio-stack itself, still nothing to do with naowalk.py.
|
||||
If TEST 2's file has clear audio: the mic and robot are fine, and the bug
|
||||
is specifically in naowalk.py's live-streaming path (network delivery,
|
||||
PyAudio device selection, etc.) - go back to the diagnostics already
|
||||
built into naowalk.py (device list / test beep / peak level) to isolate it
|
||||
further.
|
||||
|
||||
Usage:
|
||||
python3 nao_audio_diagnostic.py --ip spike.local --port 9559
|
||||
"""
|
||||
|
||||
import qi
|
||||
import argparse
|
||||
import time
|
||||
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--ip", type=str, default="127.0.0.1")
|
||||
parser.add_argument("--port", type=int, default=9559)
|
||||
args = parser.parse_args()
|
||||
|
||||
session = qi.Session()
|
||||
session.connect(f"tcp://{args.ip}:{args.port}")
|
||||
print(f"✅ Connected to {args.ip}")
|
||||
audio = session.service("ALAudioDevice")
|
||||
|
||||
# --- TEST 1: mic hardware energy, all four channels ------------------------
|
||||
print("\n=== TEST 1: Mic hardware energy (all 4 mics) ===")
|
||||
print("Make noise near the robot's head (clap, talk right up close to it) for 10s...")
|
||||
audio.enableEnergyComputation()
|
||||
front_readings, rear_readings, left_readings, right_readings = [], [], [], []
|
||||
for i in range(10):
|
||||
f, r, l, rt = (audio.getFrontMicEnergy(), audio.getRearMicEnergy(),
|
||||
audio.getLeftMicEnergy(), audio.getRightMicEnergy())
|
||||
front_readings.append(f); rear_readings.append(r); left_readings.append(l); right_readings.append(rt)
|
||||
print(f" t={i}s front={f:.2f} rear={r:.2f} left={l:.2f} right={rt:.2f}")
|
||||
time.sleep(1)
|
||||
audio.disableEnergyComputation()
|
||||
|
||||
results = {"front": front_readings, "rear": rear_readings, "left": left_readings, "right": right_readings}
|
||||
flat = [name for name, vals in results.items() if max(vals) - min(vals) < 2]
|
||||
if len(flat) == 4:
|
||||
print("⚠️ ALL FOUR mics are flat. This points to something systemic - e.g. audio "
|
||||
"inputs closed (closeAudioInputs() called somewhere and never reopened), "
|
||||
"or a low-level hardware/driver fault - rather than one damaged mic.")
|
||||
elif flat:
|
||||
print(f"⚠️ Flat mic(s): {', '.join(flat)}. The others responded normally, which "
|
||||
f"points to a fault specific to {', '.join(flat)} rather than the audio "
|
||||
f"pipeline as a whole (that pipeline is clearly working for the others).")
|
||||
else:
|
||||
print("✅ All four mics responded to noise - hardware capture is fine across the board.")
|
||||
|
||||
# --- TEST 2: local on-robot recording ---------------------------------------
|
||||
print("\n=== TEST 2: Local on-robot recording ===")
|
||||
remote_path = "/home/nao/mic_test.wav"
|
||||
print(f"Recording 5s directly to {remote_path} on the robot itself "
|
||||
f"(this bypasses the network stream and your laptop's audio stack entirely)...")
|
||||
audio.startMicrophonesRecording(remote_path)
|
||||
time.sleep(5)
|
||||
audio.stopMicrophonesRecording()
|
||||
print(f"Done. Pull it off and play it with:")
|
||||
print(f" scp nao@{args.ip}:{remote_path} .")
|
||||
print(f" (then open mic_test.wav in any media player - it's 4-channel, so check all channels)")
|
||||
print("If that file has clearly audible sound, the mic and robot are fine and the")
|
||||
print("bug is specifically in naowalk.py's live-streaming path.")
|
||||
|
||||
# --- TEST 3: robot's own speaker, via TTS (version-agnostic) ---------------
|
||||
print("\n=== TEST 3: Robot's own speaker ===")
|
||||
input("Press Enter to have the robot say a test phrase out loud...")
|
||||
session.service("ALTextToSpeech").say("Testing 1 2 3, can you hear me.")
|
||||
print("If you heard the robot speak just now, its speaker output is fine "
|
||||
"(this is unrelated to hearing its mic on your laptop, but rules out a robot-wide audio fault).")
|
||||
|
||||
print("\nDone. Report back which test(s) failed and we'll know exactly where to look.")
|
||||
+150
@@ -0,0 +1,150 @@
|
||||
# naomusic.py
|
||||
# -*- encoding: UTF-8 -*-
|
||||
|
||||
import subprocess
|
||||
import os
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
import threading
|
||||
|
||||
class NaoMusicPlayer:
|
||||
def __init__(self, session, nao_ip, ssh_port=22):
|
||||
self.session = session
|
||||
self.player = session.service("ALAudioPlayer")
|
||||
self.nao_ip = nao_ip
|
||||
self.ssh_port = ssh_port
|
||||
self.music_dir = "/home/nao/music"
|
||||
self.local_path = "/tmp/song.mp3"
|
||||
self.remote_path = f"{self.music_dir}/song.mp3"
|
||||
try:
|
||||
subprocess.run(["sshpass", "-p", "nao", "ssh", "-p", str(self.ssh_port),
|
||||
f"nao@{nao_ip}", f"mkdir -p {self.music_dir}"],
|
||||
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
|
||||
except: pass
|
||||
self.current_file = None
|
||||
self.current_task_id = None
|
||||
self.is_playing = False
|
||||
self.current_title = ""
|
||||
self.is_loading = False
|
||||
self.load_error = None
|
||||
self.status = ""
|
||||
self.progress = 0.0
|
||||
|
||||
def search(self, query):
|
||||
try:
|
||||
cmd = ["yt-dlp", f"ytsearch5:{query}", "--flat-playlist", "--dump-json",
|
||||
"--no-warnings", "--quiet"]
|
||||
result = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
|
||||
entries = []
|
||||
for line in result.stdout.strip().splitlines():
|
||||
if not line.strip():
|
||||
continue
|
||||
try:
|
||||
data = json.loads(line)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
vid = data.get("id")
|
||||
url = data.get("url")
|
||||
if not url or not url.startswith("http"):
|
||||
url = f"https://www.youtube.com/watch?v={vid}" if vid else None
|
||||
if url:
|
||||
entries.append({"url": url, "title": data.get("title", "Unknown")})
|
||||
return entries
|
||||
except Exception as e:
|
||||
print(f"Search error: {e}")
|
||||
return []
|
||||
|
||||
_DL_PROGRESS_RE = re.compile(r'\[download\]\s+([\d.]+)%')
|
||||
|
||||
def play(self, url, title):
|
||||
self.stop()
|
||||
self.current_title = title
|
||||
self.is_loading = True
|
||||
self.load_error = None
|
||||
self.status = "Starting download..."
|
||||
self.progress = 0.0
|
||||
try:
|
||||
if os.path.exists(self.local_path):
|
||||
os.remove(self.local_path)
|
||||
except: pass
|
||||
|
||||
try:
|
||||
cmd = ["yt-dlp", url,
|
||||
"-x", "--audio-format", "mp3",
|
||||
"-o", self.local_path,
|
||||
"--no-playlist", "--no-warnings", "--newline"]
|
||||
print(f"\n▶ yt-dlp {url}")
|
||||
proc = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
|
||||
text=True, bufsize=1)
|
||||
for line in proc.stdout:
|
||||
line = line.rstrip()
|
||||
if not line:
|
||||
continue
|
||||
print(f" [yt-dlp] {line}")
|
||||
m = self._DL_PROGRESS_RE.search(line)
|
||||
if m:
|
||||
self.progress = float(m.group(1))
|
||||
self.status = f"Downloading... {self.progress:.0f}%"
|
||||
elif "ExtractAudio" in line or "Extracting audio" in line:
|
||||
self.status = "Converting to mp3..."
|
||||
elif line.startswith("[download] Destination"):
|
||||
self.status = "Downloading..."
|
||||
proc.wait(timeout=120)
|
||||
|
||||
if proc.returncode != 0 or not os.path.exists(self.local_path):
|
||||
print(f"Download failed (exit {proc.returncode})")
|
||||
self.load_error = "Download failed"
|
||||
return False
|
||||
|
||||
size_mb = os.path.getsize(self.local_path) / (1024 * 1024)
|
||||
print(f"▶ scp {self.local_path} -> nao@{self.nao_ip}:{self.remote_path} ({size_mb:.1f} MB)")
|
||||
self.progress = 0.0
|
||||
upload_start = time.time()
|
||||
stop_ticker = threading.Event()
|
||||
|
||||
def ticker():
|
||||
while not stop_ticker.is_set():
|
||||
elapsed = time.time() - upload_start
|
||||
self.status = f"Uploading to NAO ({size_mb:.1f} MB)... {elapsed:.0f}s"
|
||||
stop_ticker.wait(1.0)
|
||||
|
||||
t = threading.Thread(target=ticker, daemon=True)
|
||||
t.start()
|
||||
scp_result = subprocess.run(
|
||||
["sshpass", "-p", "nao", "scp", "-P", str(self.ssh_port),
|
||||
self.local_path, f"nao@{self.nao_ip}:{self.remote_path}"],
|
||||
capture_output=True, text=True, timeout=60
|
||||
)
|
||||
stop_ticker.set()
|
||||
t.join(timeout=1.0)
|
||||
print(f" [scp] exit {scp_result.returncode}" +
|
||||
(f": {scp_result.stderr.strip()}" if scp_result.stderr.strip() else ""))
|
||||
if scp_result.returncode != 0:
|
||||
print(f"SCP failed: {scp_result.stderr.strip()}")
|
||||
self.load_error = "SCP failed"
|
||||
return False
|
||||
|
||||
self.current_file = self.remote_path
|
||||
self.status = "Playing"
|
||||
self.progress = 100.0
|
||||
self.current_task_id = self.player.playFile(self.remote_path)
|
||||
self.is_playing = True
|
||||
print(f"♪ Playing on NAO: {title} (task {self.current_task_id})")
|
||||
return True
|
||||
except Exception as e:
|
||||
print(f"Playback error: {e}")
|
||||
self.load_error = str(e)
|
||||
return False
|
||||
finally:
|
||||
self.is_loading = False
|
||||
|
||||
def stop(self):
|
||||
try:
|
||||
self.player.stopAll()
|
||||
except Exception as e:
|
||||
print(f"Stop error: {e}")
|
||||
self.is_playing = False
|
||||
self.current_file = None
|
||||
self.current_task_id = None
|
||||
self.current_title = ""
|
||||
+992
-231
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,24 @@
|
||||
# NAO Walk - Teleoperation + Music Player
|
||||
|
||||
A full-featured Python teleoperation tool for the NAO robot with live camera feed, head control, microphone streaming to your laptop, and **music playback directly on the robot**.
|
||||
|
||||
---
|
||||
|
||||
## Features
|
||||
|
||||
- **Live Camera Feed** from NAO (upscaled on PC)
|
||||
- **Keyboard Teleop** (WASD / Arrow keys)
|
||||
- **Head Control** (I J K L)
|
||||
- **Wave Gesture** (Key `1`)
|
||||
- **Live Mic Streaming** from NAO to your laptop (with adjustable gain)
|
||||
- **Music Player on NAO** — Search & play YouTube music directly through the robot's speakers
|
||||
- **TTS Chat** — Type messages for the robot to speak (`T`)
|
||||
- **Battery & Volume Display**
|
||||
|
||||
---
|
||||
|
||||
## Requirements
|
||||
|
||||
### On your laptop:
|
||||
```bash
|
||||
pip install qi opencv-python numpy pygame yt-dlp
|
||||
+13
-11
@@ -1,12 +1,14 @@
|
||||
# requirements.txt for NAO Robot Teleoperation
|
||||
# Install with: pip install -r requirements.txt
|
||||
#
|
||||
# NOTE: the NAOqi Python SDK ("qi") is NOT on PyPI. It ships separately from
|
||||
# Aldebaran/SoftBank Robotics as a platform-specific wheel (pynaoqi-python*).
|
||||
# Download it from the Softbank Robotics / Aldebaran developer portal, match
|
||||
# it to your Python version/OS, and install it with:
|
||||
# pip install ./pynaoqi-*.whl
|
||||
# See README.md for details.
|
||||
|
||||
# Core NAOqi SDK (required - install via Aldebaran/SoftBank installer)
|
||||
# qi # Usually installed with the NAOqi Python SDK
|
||||
|
||||
# Python packages
|
||||
opencv-python>=4.5.0
|
||||
pygame>=2.0.0
|
||||
numpy>=1.21.0
|
||||
|
||||
# Optional but recommended
|
||||
pillow>=9.0.0 # If you want PIL fallback for images
|
||||
opencv-python>=4.8
|
||||
numpy>=1.24
|
||||
pygame>=2.5
|
||||
pyaudio>=0.2.13
|
||||
yt-dlp>=2024.1.1
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
#!/bin/bash
|
||||
pactl set-source-volume alsa_input.pci-0000_00_1b.0.analog-stereo 32%
|
||||
#python3.11 naowalk.py --ip spike.local --port 9561 --mic-gain 9 --video-scale 3.0
|
||||
#python3.11 naowalk.py --ip 127.0.0.1 --port 9561 --mic-gain 9 --video-scale 3.0 --nao-ssh-port 2222 --media-relay 127.0.0.1:8000
|
||||
python3.11 naowalk.py --ip spike.local --port 9561 --mic-gain 9 --video-scale 3.0 --nao-ssh-port 22 --media-relay spike.local:8000
|
||||
Executable
+1
@@ -0,0 +1 @@
|
||||
ssh -N -L 9561:192.168.0.100:9561 -L 2222:192.168.0.100:22 -L 8000:192.168.0.100:8000 root@10.196.124.46
|
||||
Reference in New Issue
Block a user