I am the Melbourne Connect Chair of Digital Innovation for Society
in the School of Computing and Information Systems at the
University of Melbourne

email: tom.drummond@unimelb.edu.au


Research Topics:

Showing posts with label Robotics. Show all posts
Showing posts with label Robotics. Show all posts

Look no deeper: Recognizing places from opposing viewpoints under varying scene appearance using single-view depth estimation (with Sourav Garg, Madhu Babu, Thanuja Dharmasiri, Stephen Hausler, Niko Suenderhauf, Swagat Kumar and Michael Milford)

This paper presents an approach to solving the hard problem of finding correspondences and localisation under extreme (180°) viewpoint variations.  Depth estimation is used to filter key points for probable appearance in the opposing view and a robust descriptor is learned to aid matching.

[ICRA 2019 paper]

The Importance of Metric Learning for Robotic Vision: Open Set Recognition and Active Learning (with Ben Meyer)

This paper shows how to use metric learning for active learning.  In a metric space, examples of novel classes typically map to empty parts of the space.  This can be detected automatically using the local ratio of unlabelled to labelled densities to select examples for active labelling.

[ICRA 2019 paper]

Real-time joint semantic segmentation and depth estimation using asymmetric annotations (with Vladimir Nekrasov, Thanuja Dharmasiri, Andrew Spek, Chunhua Shen and Ian Reid)

This paper shows the benefits of simultaneously estimating semantics and depth for a monocular image input stream.  The resulting network can perform this estimate at 13ms per frame, enabling it to be used in real-time systems.

[ICRA 2019 paper]

Eng: End-to-end neural geometry for robust depth and pose estimation using cnns (with Thanuja Dharmasiri and Andrew Spek)

This paper shows how to compute camera motion using networks to estimate depth per frame and optical flow between frames with uncertainty.  These estimates and uncertainties are then combined using conventional optimisation to obtain motion.

[ACCV 2018 paper]

CReaM: Condensed real-time models for depth prediction using convolutional neural networks (with Andrew Spek and Thanuja Dharmasiri)

This paper shows how to use a complex model to train a simpler one by applying a loss to cause the embeddings in latent space to converge.  This accelerates depth estimation so that it can run at 30 frames per second on a TX2 for use in a VO/SLAM pipeline.

[IROS 2018 paper] 

Updated notes on Lie groups

The notes on Lie groups have been updated.  The theory chapter has been split into two chapters for Lie groups and projective geometry.  Sections on the Lie bracket and conics have been added.  Notes available here.

Australian Centre for Robotic Vision

The Australian Research Council have just awarded us $19M to establish a Centre of Excellence in Robotic Vision.  The centre will address some of the key challenges in enabling robots to use vision to operate in unstructured and dynamic environments alongside humans.

[ARC Announcement]

Multiview Image Compression and Transmission Techniques in Wireless Multimedia Sensor Networks: A Survey (with Max Wang and Ahmet Sekercioglu)

This paper presents a survey of recent research works on multiview image compression and transmission techniques developed for Wireless Multimedia Sensor Networks (WMSNs). We classify them into two categories with respect to the coding methods adopted: (i) in-network processing with joint coding schemes, and (ii) distributed source coding schemes. The survey also includes a comprehensive evaluation of the limitations of each approach.

[ICDSC 2013 paper]

A Real-Time Distributed Relative Pose Estimation Algorithm for RGB-D Camera Equipped Visual Sensor Networks (with Max Wang and Ahmet Sekercioglu)

In this paper, we present a distributed, peer-to-peer algorithm for relative pose estimation in a network of mobile robots equipped with RGB-D cameras acting as a visual sensor network.  Our algorithm uses the depth information to estimate the relative pose of a robot when camera sensors mounted on different robots observe a common scene from different angles of view.

[ICDSC 2013 paper]

Robust egomotion estimation using ICP in inverse depth coordinates (with Dennis Lui, Titus Tang and Wai Ho Li)

This paper presents a 6 degrees of freedom egomotion estimation method using Iterative Closest Point (ICP) for low cost and low accuracy range cameras. Instead of Euclidean coordinates, the method uses inverse depth coordinates which better conforms to the error characteristics of raw sensor data. Extensive experiments were performed to evaluate different combinations of error metrics and parameters. The result is a real-time system that is accurate and robust across a variety of motion trajectories.


Visual localisation of a robot with an external RGBD sensor (with Winston Yii, Nalika Damayanthi and Wai Ho Li)

This paper presents a novel approach to visual localisation that uses a camera on the robot coupled wirelessly to an external RGB-D sensor. Unlike systems where an external sensor observes the robot, our approach merely assumes the robots camera and external sensor share a portion of their field of view. Experiments were performed using a Microsoft Kinect as the external sensor and a small mobile robot. The robot carries a smartphone, which acts as its camera, sensor processor, control platform and wireless link. Computational effort is distributed between the smartphone and a host PC connected to the Kinect. Experimental results show that the approach is accurate and robust in dynamic environments with substantial object movement and occlusions.  This work won the best student paper prize at ACRA 2011.
[ACRA 2011 paper]

eBug - an open robotics platform for teaching and research (with Nick D'Ademo, Dennis Lui, Wai Ho Li and Ahmet Sekercioglu)


The eBug is a low-cost and open robotics platform designed for undergraduate teaching and academic research in areas such as multimedia smart sensor networks, distributed control, mobile wireless communication algorithms and swarm robotics. The platform is easy to use, modular and extensible.


[ACRA 2011 paper]





Visually guided aerial robotics (with Chris Kemp)

This quad-rotor helicopter carries a miniature transmitting video camera. The video stream is received and fed into a computer which uses a visual tracking algorithm to compute the helicopter's position and generate control signals. These are sent to the remote control unit to stabilise the helicopter. The inset in the top right of the video shows the computer's view from the on-board camera. This work uses the dynamic measurement clustering technique described below. 

[Video in corridor] [Video in lab] [Video showing response to disturbance] [Chris Kemp's PhD Thesis]


Edge-based tracking

This was our first research into edge-based tracking based on the Harris' RAPiD tracker. The videos below were made in 1999 on a 225 MHz SGI. This work won the Industry prize at BMVC 1999

[Prizewinning 1999 BMVC Paper] [Video1] [Video2 (with occlusion)]