# Triggering Demos with Voice Commands

**URL:** <https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561>\
**Category:** Ask\
**Created:** [January 12, 2023, 4:27pm UTC](https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561 "2023-01-12T16:27:57Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jgangemi](https://avatars.discourse-cdn.com/v4/letter/j/8491ac/32.png) [@jgangemi](https://forum.hello-robot.com/u/jgangemi)\
**Post date:** [January 12, 2023, 4:27pm UTC](https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561/1 "2023-01-12T16:27:57Z")

</div>

Hello!

I am trying to figure out how to get stretch to be controlled by voice commands. I figured a good starting place would be to trigger one of the demos (like grasp\_object) to initiate by saying a certain word or phrase rather than a keyboard input. I have already played around with the stretch demos, and have gotten the respeaker speech to text functions working. I am new to hello robot and ROS and not sure I quite understand where to start in putting this all together now. Has anyone tried/ had any success communicating with stretch through voice prompts yet?

Thanks,  
Julia

---

<div class="post-metadata">

**Author:** ![lamsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/lamsey/32/177_2.png) [@lamsey](https://forum.hello-robot.com/u/lamsey)\
**Post date:** [January 12, 2023, 7:32pm UTC](https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561/2 "2023-01-12T19:32:29Z")

</div>

Hello! I’ve had success with using voice prompts on Stretch + ROS. Here is a [link](https://github.com/mlamsey/stretch-caregiving-class/blob/main/nodes/stretch_with_stretch/color_listener.py) to a ROS node that records a snippet of audio using Stretch’s internal mic, and counts how many unique colors were spoken in the snippet. For your application, you could create logic to execute different actions based on the strings detected in the recording. You could use the [roslaunch API](http://wiki.ros.org/roslaunch/API%20Usage) to launch a launch file from `stretch_demos`, but make sure you only run one Stretch driver at a time.

We used the `SpeechRecognition` [python package](https://pypi.org/project/SpeechRecognition/). PyPI docs had everything we needed to set it up, such as this [example](https://github.com/Uberi/speech_recognition/blob/master/examples/microphone_recognition.py).

See this [launch file](https://github.com/mlamsey/stretch-caregiving-class/blob/main/launch/stretch_with_stretch.launch) for how we included the node in our application.

Snippet:

```auto
import speech_recognition as sr

recording_length = 10. # seconds

r = sr.Recognizer()
with sr.Microphone() as source:
    # record
    audio_clip = r.record(source, duration=recording_length)

    # recognize
    try:
        text_string = r.recognize_google(audio_clip) # see PyPI docs for other options
    except sr.UnknownValueError:
        print("Speech recognizer could not understand audio")
    except sr.RequestError as e:
        print("Speech recognition error; {0}".format(e))

    # execute
    if "grasp" in text_string:
        roslaunch_grasp_demo() # implement this elsewhere using the roslaunch package
    else:
        print("No command found in string")

```

---

<div class="post-metadata">

**Author:** ![bshah](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/bshah/32/55_2.png) [@bshah](https://forum.hello-robot.com/u/bshah)\
**Post date:** [January 12, 2023, 8:52pm UTC](https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561/3 "2023-01-12T20:52:53Z")

</div>

Hi @jgangemi, welcome to the forum! It’s a great question, and I really like @lamsey’s reply. I would approach it pretty much the way that he has described.

I wanted to highlight a few other implementations I’ve seen from members on this forum, and hopefully, they will be useful for you.

- @asanchez made a tutorial that is available on our docs, called [Voice Teleoperation of Base](https://docs.hello-robot.com/0.2/stretch-tutorials/ros1/example_9/). Phrases like “forward” triggers motions of the mobile base, and there a code explanation section that breaks down what each part of the code is doing.
- @hello-garv created a node for the Stretch Web Interface, called [speech\_commands.py](https://github.com/hcrlab/stretch_web_interface/blob/master/nodes/speech_commands.py), which builds on Alan’s tutorial. It enables voice teleop for Stretch’s other joints, but also triggers ROS services to replay saved poses. Her [post](https://forum.hello-robot.com/t/navigation-with-aruco-tags/483) has more details.
- @FergusKidd [posted](https://forum.hello-robot.com/t/voice-control-as-input/44/5) about example code using Azure speech-to-text and LUIS (a cloud service that understands intent from phrases) to control a Stretch. Their project went even further to use text-to-speech and Q&A services to get Stretch to give answers back.

---

<div class="post-metadata">

**Author:** ![chintujaguar](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/chintujaguar/32/260_2.png) [@chintujaguar](https://forum.hello-robot.com/u/chintujaguar)\
**Post date:** [January 13, 2023, 12:32am UTC](https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561/4 "2023-01-13T00:32:40Z")

</div>

Hey @jgangemi,

To trigger the grasp object demo using speech you could build up on @lamsey’s response while using the [keyboard\_teleop](https://github.com/hello-robot/stretch_ros/blob/master/stretch_core/nodes/keyboard_teleop) node as a reference. This node illustrates the way the ‘/grasp\_object/trigger\_grasp\_object’ service can be triggered using ROS ServiceProxy like [here](https://github.com/hello-robot/stretch_ros/blob/master/stretch_core/nodes/keyboard_teleop#L207). Adding to @lamsey’s response, here’s a snippet that should work:

```python
recording_length = 10. # seconds

r = sr.Recognizer()

rospy.wait_for_service('/grasp_object/trigger_grasp_object')
rospy.loginfo('Node connected to /grasp_object/trigger_grasp_object.')
trigger_grasp_object_service = rospy.ServiceProxy('/grasp_object/trigger_grasp_object', Trigger)

with sr.Microphone() as source:
    # record
    audio_clip = r.record(source, duration=recording_length)

    # recognize
    try:
        text_string = r.recognize_google(audio_clip) # see PyPI docs for other options
    except sr.UnknownValueError:
        print("Speech recognizer could not understand audio")
    except sr.RequestError as e:
        print("Speech recognition error; {0}".format(e))

    # execute
    if "grasp" in text_string:
        trigger_request = TriggerRequest() 
        trigger_result = trigger_grasp_object_service(trigger_request)
        print('trigger_result = {0}'.format(trigger_result))
    else:
        print("No command found in string")

```

Once you wrap this in a ROS node and launch it along with the grasp\_object node like in [this](https://github.com/hello-robot/stretch_ros/blob/master/stretch_demos/launch/grasp_object.launch) launch file, you’d have a working grasp object demo that runs using voice commands.

Let me know if you need further assistance with this.

Best,  
Chintan

---

<div class="post-metadata">

**Author:** ![jgangemi](https://avatars.discourse-cdn.com/v4/letter/j/8491ac/32.png) [@jgangemi](https://forum.hello-robot.com/u/jgangemi)\
**Post date:** [January 13, 2023, 3:11pm UTC](https://forum.hello-robot.com/t/triggering-demos-with-voice-commands/561/5 "2023-01-13T15:11:06Z")

</div>

Thank you all so much!! These are all very helpful examples, and I will be trying them all out today. I will let you know how it goes 🙂 Thank you all again!!
