Evaluating LLM-Based Agent Orchestration Frameworks for Industrial Deployment

Copy Link

LLM-based agent orchestration frameworks enable the coordination of multiple AI agents to complete complex, multi-step tasks that exceed the capabilities of a single model. Tools such as LangGraph, AutoGen, and CrewAI have moved rapidly from research prototypes to production deployments across finance, software engineering, and operations. However, the academic literature has not kept pace with this adoption — there is little systematic understanding of how these frameworks behave under production conditions, how organisations customise them, or what architectural decisions lead to reliable versus brittle deployments. This project takes an empirical, industry-grounded approach to understanding the lifecycle of orchestrated agent systems: from framework selection and architecture design through to monitoring, failure recovery, and iteration. The goal is to produce findings that are both theoretically grounded and immediately applicable to practitioners deploying these systems at scale.

Aim  

This project aims to develop a rigorous empirical understanding of how LLM-based agent orchestration frameworks are being used in industrial practice and to identify the key factors that determine their reliability, scalability, and fitness for purpose in real-world deployments. Specifically, the project aims to: (1) characterise the landscape of current orchestration framework adoption across industry sectors; (2) identify the common architectural patterns and engineering decisions that practitioners employ when deploying multi-agent systems; (3) evaluate the performance and failure characteristics of leading frameworks (LangGraph, AutoGen, CrewAI) under realistic production conditions; and (4) develop a validated framework of design principles and evaluation criteria to guide industry practitioners and inform future research on agent system engineering.

Objectives 

The project will pursue the following concrete objectives: (1) conduct a systematic review of LLM-based orchestration frameworks, covering architectural design, supported agent patterns, and documented industrial use cases; (2) carry out structured interviews and surveys with industry practitioners across sectors actively deploying multi-agent systems; (3) design and execute controlled benchmarking experiments comparing LangGraph, AutoGen, and at least one additional framework across task complexity, latency, cost, and failure recovery dimensions; (4) develop and validate a taxonomy of failure modes specific to orchestrated LLM agent systems in production; (5) co-develop, through iterative engagement with industry partners, a set of evidence-based design guidelines for robust multi-agent deployment; and (6) disseminate findings through peer-reviewed publications and an open-access evaluation toolkit for practitioners.

Significance 

Multi-agent LLM systems are being deployed in high-stakes industrial contexts — including financial analysis, software development automation, and supply chain management — yet the engineering discipline to support reliable deployment remains immature. This project addresses a critical gap: the absence of empirically validated guidance for practitioners navigating framework selection, architecture design, and production operations. By producing the first large-scale empirical study of LLM orchestration frameworks in industry, this research will directly inform how organisations invest in and govern AI automation infrastructure. It will also contribute foundational knowledge to the emerging discipline of agent system engineering, establishing evaluation standards and failure taxonomies that can anchor future academic work. At a broader level, the project will help ensure that the rapid industrialisation of multi-agent AI proceeds on a more principled and evidence-based footing.

Ideal Candidate 

Applicants should hold a strong undergraduate or honours degree in computer science, software engineering, or a closely related discipline. Demonstrated programming proficiency in Python is essential, with experience in LLM APIs (e.g. OpenAI, Anthropic) or agent frameworks highly regarded. Familiarity with research methods — including literature review, experimental design, and qualitative data collection — is desirable. Strong written communication skills are important given the publication expectations of the role. An interest in bridging academic rigour with industry practice is important. Additionally, the applicants should meet the eligibility criteria for entry into a PhD program at Curtin University. 

This project is open to International and Domestic applicants. 

Scholarship  

If you are identified as the preferred candidate for this project, you may be considered for an RTP scholarship

Enquires and How to Apply 

For enquires about this opportunity contact Associate Professor Susannah Soon at Susannah.soon@curtin.edu.au

To formally apply submit an Expression of Interest to Associate Professor Susannah Soon during the Central Scholarship round (July 1st – July 31st 2026) 

Copy Link