Using Machine Learning to develop relevant acoustic and kinematic metrics of speech

Copy Link

Most children with Developmental Delay experience speech impairment. Thus, early identification of speech and language impairment potentially represents the point of entry to intervention services. Timely and accurate diagnosis followed by the right intervention mitigates the short- and long-term negative consequences of Developmental Delay. However, current clinical practice relies on lengthy and subjective measures that are vulnerable to bias, leading to potential misdiagnosis, delaying the provision of the right intervention. There is an urgent need to replace these practices with efficient and objective diagnostic tools that can also be administered remotely.
Current assessment practices for speech impairment comprise administration of a (paper-based) standardised norm- or criterion-referenced speech assessment, by a speech-language pathologist (S-LP). Following the face-to-face test administration, S-LPs will typically undertake auditory-perceptual analyses for diagnosis and determination of evidence-based treatment pathways. However, recent data reveals S-LPs report lack of confidence and competence in differential diagnosis of subtypes of speech impairments.
We are developing novel algorithms for the Speech-Movement and Acoustic Analysis Tracking (SMAAT) —to objectively identify and measure biomarkers of speech impairment to assist in diagnosis.

Aim  

Advancements in technology and artificial intelligence have been used to support improvements in healthcare efficiencies and outcomes for participants. In particular, computer aided systems have been developed to improve the accuracy of diagnostics and support the interpretation of analyse [12]. The importance of these advances has become particularly apparent with the COVID-19 pandemic [13]. Yet, the utilisation of artificial intelligence (AI) techniques to support diagnosis by S-LPs remains largely under-utilised [14]. To date, most applications “simply provide an alternate method of stimulus presentation” [14]. This therefore represents a research-to-practice gap and indicates the need for new innovations that include the development of suitably sensitive automated speech assessment tools, with robust and documented psychometric properties.

Our team focuses on the extraction and validation of clinically relevant acoustic and kinematic metrics of speech, thus providing an innovative solution to an unmet need. Our research uses AI to develop algorithms to provide an accurate automated solution for diagnosis of speech impairment.

Objectives 

The objective scores are currently derived from 28 discrete facial landmarks, based on distinct facial features and limited to video sequences captured by a single camera. While having discrete information allows to derive kinematic profiles and to classify between groups with and without speech impairment it does not provide a complete map of the face. As a result, facial characteristics not contained in the 28 extracted landmarks are excluded from analysis.

As an alternative to facial landmarks, the aim of this project is to implement a mesh representing the whole face as a continuous surface. The use of a mesh could allow us to quantify facial movements currently not available using landmark detection, and explore possible facial biomarkers for the diagnosis of speech disorders. For instance, hyperactivity of the mentalis muscle and lower lip may interfere with lip movements for labio-dental sounds.

  1. Design a workflow to derive full 3D facial kinematic meshes based on the existing camera set up used during data capture.
  2. Extract and validate low-cost structured-light and stereo-photogrammetry technology to derive 3D facial kinematic meshes.
  3. Exploring facial features in the full 3D facial surface data which have not been able to be considered in the current approaches for the diagnoses of speech impairment.

Significance 

Speech Sound Disorders (SSD) are common in children and account for 75% of S-LP caseloads. SSD is an umbrella term referring to any difficulty or combination of difficulties with speech planning and production and there are ~46 different interventions. Without access to the right intervention at the right time, children with SSD experience social and educational disadvantages, activity limitations, frustration, and isolation that contribute to long-term societal health and economic costs. The results address critical issues in the assessment practices of SLPs. The research outputs of this study are significant and expected to contribute to clinical practice change, with the potential to: (1) increase the timeliness, efficiency, and accuracy of diagnosis of SSD for children with development delay; and (2) reduce the time and financial cost of test administration in the determination of SSD. Our work addresses a significant evidence-practice gap with outcomes directed at impacting knowledge paradigms (building capacity and informed decision making), practice and health systems.

Ideal Candidate 

We are looking for a self-motivated PhD candidate with excellent organisation, problem-solving and project management skills.

Background/Skills in required for the project includes:

  • Background in photogrammetry and/or applied mathematics, including computer science
  • Programming experience (such as matlab, C++ or pyton),

The study will be interdisciplinary, and a willingness to work with experts from Allied Health is essential. Additionally, the applicants should meet the eligibility criteria for entry into a PhD program at Curtin University. 

This project is open to International and Domestic applicants. 

Scholarship  

If you are identified as the preferred candidate for this project, you may be considered for an RTP scholarship

Enquires and How to Apply 

For enquires about this opportunity contact Associate Professor Petra Helmholz at Petra.Helmholz@curtin.edu.au

To formally apply submit an Expression of Interest to Associate Professor Petra Helmholz during the Central Scholarship round (July 1st – July 31st 2026) 

Copy Link