# Automatically grabbing objects with Stretch3

**URL:** <https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024>\
**Category:** Ask\
**Created:** [June 28, 2024, 3:57am UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024 "2024-06-28T03:57:08Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![allen](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/allen/32/914_2.png) [@allen](https://forum.hello-robot.com/u/allen)\
**Post date:** [June 28, 2024, 3:57am UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/1 "2024-06-28T03:57:08Z")

</div>

Hello,

I’ve been working on programming Stretch 3 to detect specific objects, like bottles. I’ve explored stretch\_visual\_servoing but couldn’t find any configuration options for specifying the object model name. Could someone guide me on how to set this up? Any assistance would be greatly appreciated.

Thank you!

---

<div class="post-metadata">

**Author:** ![Mohamed\_Fazil](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/mohamed_fazil/32/261_2.png) [@Mohamed\_Fazil](https://forum.hello-robot.com/u/Mohamed_Fazil)\
**Post date:** [June 28, 2024, 2:45pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/2 "2024-06-28T14:45:50Z")

</div>

Hey @allen, The Visual Servoing demo available in stretch\_visual\_servoing should be able to find the objects in the [Yolov8 segmentation model](https://docs.ultralytics.com/tasks/segment/) classes. Currently the demo is hard coded to grasp only `apple` and `sports ball`, but you can modified it or add objects to grasp by modifying [yolo\_servo\_perception.py L#83](https://github.com/hello-robot/stretch_visual_servoing/blob/main/yolo_servo_perception.py#L83).

---

<div class="post-metadata">

**Author:** ![allen](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/allen/32/914_2.png) [@allen](https://forum.hello-robot.com/u/allen)\
**Post date:** [June 28, 2024, 9:01pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/3 "2024-06-28T21:01:03Z")

</div>

Hi @Mohamed_Fazil,

Thank you for your response. I followed your approach and added ‘bottle’ to the class name list: `if class_name in ['apple', 'sports ball', 'bottle']:`

However, when I ran `python3 visual_servoing_demo.py` again, the camera still did not detect the bottle. I also tested this with a tennis ball, but it didn’t work either.

---

<div class="post-metadata">

**Author:** ![hello-lamsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/hello-lamsey/32/740_2.png) [@hello-lamsey](https://forum.hello-robot.com/u/hello-lamsey)\
**Post date:** [July 1, 2024, 8:04pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/4 "2024-07-01T20:04:44Z")

</div>

Hi Allen,

I just freshly cloned + gave the visual servoing demo a try, and here are my results. I added an orange, a fork, and a toothbrush to the YOLO object list. Let me know if following these steps solves your troubles.

# 1. Add objects to YOLO detector

I changed line 83 in [yolo\_servo\_perception.py](https://github.com/hello-robot/stretch_visual_servoing/blob/main/yolo_servo_perception.py#L83) to:

`if class_name in ['apple', 'sports ball', 'orange', 'fork', 'toothbrush']:`

You could add `bottle` to this list. I found a full list of YOLO classes [here](https://stackoverflow.com/questions/77477793/class-ids-and-their-relevant-class-names-for-yolov8-model).

# 2. Test streaming from camera and running YOLO

In two separate terminals, I ran:

- `python3 send_d405_images.py `
- `python3 recv_and_yolo_d405_images.py`

Then, I e-stopped the robot and moved its arm over the table. This resulted in the following visualizations. Note how the toothbrush is not recognized - YOLO had a tough time with this object. If you can’t see any objects detected at this stage, there may be other issues with YOLO.

 ![Screenshot from 2024-07-01 15-28-24](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/8355f716ec34af55a6ee7a3303acbf33f8cb5c7c.png)  
Above: orange

 ![Screenshot from 2024-07-01 15-27-52](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/a19713c08f692cd916da922c83d624459c231343.png)  
Above: fork

 ![Screenshot from 2024-07-01 15-29-02](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/97d51041b5eed7bf7a49bde657881a10e3f362e4.png)  
Above: unrecognized toothbrush

# 3. Run visual servoing

To run the visual servoing with YOLO, the following three processes should be running simultaneously in separate terminals:

- `python3 send_d405_images.py `
- `python3 recv_and_yolo_d405_images.py`
- `python3 visual_servoing_demo.py -y` (the `-y` flag enables YOLO)

After running the code, I e-stopped the robot, pointed the wrist at the table, and turned the e-stop off. Then, I placed each object on the table. Below is what that looked like for me. Note how the servoing recovers after bumping the orange backwards.

![servo](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/d7deb3514a6afd879c520ff99e5685c6a993710b.gif)  
Above: visual servoing in action!

I hope that this helps you debug your issues!

---

<div class="post-metadata">

**Author:** ![allen](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/allen/32/914_2.png) [@allen](https://forum.hello-robot.com/u/allen)\
**Post date:** [July 1, 2024, 9:08pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/5 "2024-07-01T21:08:43Z")

</div>

Hi @hello-lamsey !

Thank you so much for your response! It was very helpful and now stretch3 is able to detect my bottle.  
Nevertheless, the issue I encounter right now is that the robot was not able to calculate the distance between the object and the gripper correctly result in grabbing air in front of the object like in the image below:

 ![Photo on 7-1-24 at 4.07 PM](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/a26dd7e23d8a4da7b442fcf230ccf3565ed3607e.jpeg)  
Could you point me into the direction on how to fix this issue?

Thansk!

---

<div class="post-metadata">

**Author:** ![hello-lamsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/hello-lamsey/32/740_2.png) [@hello-lamsey](https://forum.hello-robot.com/u/hello-lamsey)\
**Post date:** [July 2, 2024, 5:35pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/6 "2024-07-02T17:35:20Z")

</div>

Hi @allen,

I’m glad that you got the code running!

The first thing to check would be that the visual servoing code is calculating distances correctly based on the camera feed. The coordinate system for the image (and object position) are shown below. The `[x,y,z]` position (cm) of the center of the object is printed over the object.

 ![vis_cs](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/ea856fcaddc2b681b72613515bca833f691f379c.png)  
Above: image coordinate system (position approximate). Positive Z (blue) is pointing into the screen (or directly out of the camera).

To achieve grasping of the tennis ball, my system reports a tennis ball position of approximately `[x,y,z] = [0, 4, 17]` centimeters, which matches the [default grasp position](https://github.com/hello-robot/stretch_visual_servoing/blob/main/visual_servoing_demo.py#L115) defined in the demo.

 ![grasp_pos](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/552b36adb1b424c195c378bb77dd23d7e58c6f4d.jpeg)  
Above: ball in grasp location.

Can you share a screenshot / screen recording of the YOLO render window to debug measurements? It is possible that the `z` distance is not being measured correctly.

Also note: the target grasp position between the robot’s fingertips is also [calculated online](https://github.com/hello-robot/stretch_visual_servoing/blob/main/visual_servoing_demo.py#L477), so the target grasp position may vary between robots.

---

<div class="post-metadata">

**Author:** ![allen](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/allen/32/914_2.png) [@allen](https://forum.hello-robot.com/u/allen)\
**Post date:** [July 2, 2024, 7:44pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/7 "2024-07-02T19:44:50Z")

</div>

Hi @hello-lamsey

Thanks again for your reply. After testing out, the robot seems not be able to nivagate to the proper position for grapping. As shown below, the position the robot started grabbing is `1.6,4.2,12.9`, also it seems like the estimated position of the grippers are different than yours too. Furthermore, when the robot move it’s gripper to this position, it stops and starts grabbing and results in grabbing nothing.

 ![Screenshot 2024-07-02 at 2.38.43 PM](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/ffc7df727e0bd7edfa5e26f37b04c3ca15c4acf1.jpeg)  
 ![Photo on 7-2-24 at 2.44 PM](https://canada1.discourse-cdn.com/flex030/uploads/hello_robot2/original/1X/ed06adf5aa851c932a9e993796ab50565148d9b4.jpeg)

---

<div class="post-metadata">

**Author:** ![hello-lamsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/hello-lamsey/32/740_2.png) [@hello-lamsey](https://forum.hello-robot.com/u/hello-lamsey)\
**Post date:** [July 3, 2024, 9:31pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/8 "2024-07-03T21:31:42Z")

</div>

Hi @allen,

It looks like the depth readings for the bottle are off. The model thinks that the bottle is ~3cm wide and close to the camera, which does not appear to match reality. I tried with a transparent bottle on my end, and while YOLO did not always see the bottle, the depth readings appeared correct. Do you observe erroneous distance / size readings when looking at an opaque bottle as well?

I see from an earlier post in the thread that you were having trouble with the tennis ball as well. Have you tested whether the visual servoing demo works for the tennis ball and / or the ArUco cube (running without YOLO)? If so, what were the results?

Also, it is odd that the grasp begins when the object’s `z=12.9`. While running visual servoing, I see values such as `'grasp_center_xyz': array([0.00646052, 0.06986317, 0.17505255])` in the YOLO results printout in the visual servoing terminal. What value are you seeing for `'grasp_center_xyz'` in your terminal output?

Example YOLO results output:  
`yolo_results = {'fingertips': {'right': {'pos': array([0.09063274, 0.03583154, 0.1768219]), 'x_axis': array([-0.68398136, 0.03269413, -0.72876649]), 'y_axis': array([0.08904867, -0.98778254, -0.12789051]), 'z_axis': array([-0.72404409, -0.15237041, 0.67271347])}, 'left': {'pos': array([-0.07698445, 0.03379969, 0.17573048]), 'x_axis': array([0.69878856, 0.10996119, -0.70682607]), 'y_axis': array([-0.0210888 , 0.99085159, 0.13329814]), 'z_axis': array([0.71501735, -0.0782411 , 0.6947147])}}, 'yolo': [{'name': 'orange', 'confidence': 0.29331216, 'width_m': 0.04629309170714137, 'estimated_z_m': 0.15356746551650566, 'grasp_center_xyz': array([0.00646052, 0.06986317, 0.17505255]), 'left_side_xyz': array([-0.01747896, 0.06128851, 0.15356747]), 'right_side_xyz': array([0.02881413, 0.06128851, 0.15356747])}]}`

---

<div class="post-metadata">

**Author:** ![allen](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/allen/32/914_2.png) [@allen](https://forum.hello-robot.com/u/allen)\
**Post date:** [July 5, 2024, 9:06pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/9 "2024-07-05T21:06:59Z")

</div>

Hi @hello-lamsey ,

I have included the links to the recordings of the results for grabbing the transparent and opaque bottles. It seems like the robot struggles a little bit but eventually manages to grab the opaque bottle, whereas it still cannot grab the transparent bottle. Although I don’t have a video for grabbing the ArUco cube, it works perfectly fine.

[Transparent bottle video](https://drive.google.com/file/d/1awXCo2qbtU_aIep0O0XmIS049nGQ6Kf0/view?usp=sharing)

[Opaque bottle video](https://drive.google.com/file/d/1QmujJ9hZQ_ohHdcWG_lwYTuoUFvKEBfK/view?usp=sharing)

For the ‘grasp\_center\_xyz’ terminal output when grabbing the transparent bottle, I am seeing values such as:

- `'grasp_center_xyz': array([0.01307648, -0.03386912, 0.23353366])`
- `'grasp_center_xyz': array([0.01090819, -0.0208785 , 0.23958624])`
- `'grasp_center_xyz': array([0.01349765, 0.0044698 , 0.26781633])`
- `'grasp_center_xyz': array([0.01890173, 0.06022887, 0.23040281])`
- `'grasp_center_xyz': array([0.01301666, 0.05514316, 0.22045396])`
- `'grasp_center_xyz': array([0.01462979, 0.03248335, 0.17890907])`

These are quite different from the values you are seeing.

Best regards,  
Allen

---

<div class="post-metadata">

**Author:** ![hello-lamsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/hello-lamsey/32/740_2.png) [@hello-lamsey](https://forum.hello-robot.com/u/hello-lamsey)\
**Post date:** [July 9, 2024, 6:58pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/10 "2024-07-09T18:58:18Z")

</div>

Hi @allen,

Glad to hear that grabbing the ArUco cube works well! Also, those grasp centers do seem to vary a decent amount. It might be within tolerances for the code to successfully grasp some objects, but if performance is poor, then factors such as the lighting conditions may be negatively affecting the estimate of the fingertip ArUco markers.

Regarding the transparent bottle, the [d405 wrist camera](https://www.intelrealsense.com/depth-camera-d405/) uses RGB stereo image pairs to compute depth. It may struggle with transparent objects such as the water bottle that you are using; I recommend trying to use opaque objects when possible.

The original visual servoing code was tuned to work with the cube and the tennis ball, so grasping other objects reliably may take some tweaking of parameters. If you are interested in improving performance while grasping the opaque bottle, here are some code bits that could be adjusted:

- Change the `grasp_depth` in [yolo\_servo\_perception.py](https://github.com/hello-robot/stretch_visual_servoing/blob/main/yolo_servo_perception.py#L157): see `YoloServoPerception.apply()`

- Change (reduce) the speeds for the arm and lift in the `retract` state in [visual\_servoing\_demo.py](https://github.com/hello-robot/stretch_visual_servoing/blob/main/visual_servoing_demo.py#L532): see the `behavior == 'retract'` block inside `main()` and adjust the `cmd` dictionary.

Since this code uses velocity control for the robot’s joints, **be careful not to turn the speeds up too high!** I also recommend staying near the E-Stop in case the parameters that you adjust cause the robot to move in an undesirable way.

Last, if you are interested in an alternative approach to visual servoing, consider checking out the [stretch\_forcesight](https://github.com/hello-robot/stretch_forcesight) repository. This is an **experimental** deep learning-based approach to planning force and position targets for Stretch’s end effector to achieve while performing semantically labeled tasks, such as picking up a cup. Note that the deep model in `stretch_forcesight` was trained using an experimental camera mount on the end effector, as well as older models of Stretch, so it will likely not perform as well as described in the [original paper](https://force-sight.github.io/) out-of-the-box.

---

<div class="post-metadata">

**Author:** ![allen](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/allen/32/914_2.png) [@allen](https://forum.hello-robot.com/u/allen)\
**Post date:** [July 10, 2024, 9:55pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/11 "2024-07-10T21:55:22Z")

</div>

**Hi @hello-lamsey!**

Thanks for the additional information. I’m also curious if there is a ROS 2 package with similar functionality to _stretch\_visual\_servoing_ or _stretch\_forcesight_. Specifically, I’m looking for a ROS 2 package or tool that can provide a message indicating whether the robot has successfully grabbed an object or not.

Do you know of any existing packages or solutions that offer this capability?

Thanks in advance for your help!

---

<div class="post-metadata">

**Author:** ![hello-lamsey](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/hello-lamsey/32/740_2.png) [@hello-lamsey](https://forum.hello-robot.com/u/hello-lamsey)\
**Post date:** [July 15, 2024, 7:03pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/12 "2024-07-15T19:03:54Z")

</div>

Hi @allen,

Currently, there is not an official ROS2 package for visual servoing (or forcesight) with Stretch. I have opened an [issue](https://github.com/hello-robot/stretch_visual_servoing/issues/6) on github that requests the creation of a ROS2 package for visual servoing.

One current limitation is that the ROS2 `stretch_driver` does not directly expose a velocity controller for the robot’s joints yet. Velocity control is used in the pythonic demo and is a key enabler of fast, dexterous tracking. Running velocity control over ROS2 may present issues related to factors like communication latency, so this could be a bit tricky to implement safely as a prerequisite for a visual servoing package.

If there is general interest in the creation of a visual servoing ROS2 package, let us know!

---

<div class="post-metadata">

**Author:** ![lstegner](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/lstegner/32/90_2.png) [@lstegner](https://forum.hello-robot.com/u/lstegner)\
**Post date:** [August 15, 2024, 9:41pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/13 "2024-08-15T21:41:35Z")

</div>

Hi @hello-lamsey,  
Thank you for the response! Allen is working with me on Stretch development–we ultimately need to be able to switch between doing some navigation (we are using the nav2 package, and it works for us) and doing some manipulation. Given that there seems to be a block with easily porting the visual servoing to a ROS2 package, do you have any recommendation for another approach we could take?

Allen was also exploring switching between the nav2 calls and the visual servoing script, but naturally when the visual servoing script ends and switches back to navigation, the arm does not continue maintaining its position holding the item.

---

<div class="post-metadata">

**Author:** ![bshah](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.hello-robot.com/bshah/32/55_2.png) [@bshah](https://forum.hello-robot.com/u/bshah)\
**Post date:** [September 3, 2024, 8:11pm UTC](https://forum.hello-robot.com/t/automatically-grabbing-objects-with-stretch3/1024/14 "2024-09-03T20:11:32Z")

</div>

Hi @lstegner, recently, the ROS2 driver has gotten a new [“streaming” mode](https://forum.hello-robot.com/t/software-drop-august-23-2024/1092#p-2837-streaming-mode-in-ros2-driver-2), which should make it much simpler to port the visual servoing code to ROS2. If you’re interested in this, I recommend reaching out to @cpaxton, since he has been experimenting with this new mode to run visual servoing code.
