In the world of generative AI, what does good even look like

One of the most frequently cited challenges organisations face when implementing probabilistic and agentic AI is determining a true return on investment. But one of the hardest questions behind that challenge is beguilingly simple. If an organisation is looking for a good outcome, how does it determine what “good” actually is?

This was perhaps the most important question posed at the recent Sydney AI Engineering and Infrastructure Summit, hosted by Clutch Events, and it was one that drew little consensus among the speakers.

Determining “good” in the era of deterministic systems was relatively straightforward. Performance could usually be measured in terms such as accuracy, speed, reliability, and cost.

But many of the challenges to which probabilistic and agentic AI systems are now being applied defy such simple measures.

A chatbot engaging with a human being might communicate more efficiently than a person, but that does not mean the experience is better. Organisations can measure proxies, such as whether a conversation leads to a desired action, but that bulk metric says little about the quality of the interaction itself. Sentiment surveys provide another approximation, but they are limited by who responds and by the inherently subjective nature of the feedback.

Determining what “good” looks like becomes an interesting new challenge for AI infrastructure and engineering professionals. It is not something they needed to think about to the same extent in the earlier generation of technology, where performance was more straightforward to provision, measure, and monitor.

It also points to a growing need for engineers and business leaders to work more closely together. The technical system cannot be separated from the use case it serves. Organisations need to connect what the AI is doing to the outcome the business is trying to achieve, and then determine whether that outcome represents genuine value.

And defining value is only one of the challenges organisations encounter as they try to move AI from experimentation into production.

Another challenge discussed throughout the day was skills availability, made more difficult by the reluctance of many organisations to invest sufficiently in the training and upskilling required to build those capabilities internally.

The challenge of scaling AI becomes more complicated again when organisations consider where their models should run. The default approach has been to consume frontier models through the public cloud. But rising costs are prompting some organisations to explore running open-weight models in their own environments or through third-party infrastructure providers.

The result is increasingly a hybrid model, along with the accompanying headache of constantly managing different cost structures. Running models outside the public cloud can reduce some costs and give organisations greater control, but it also introduces trade-offs. Security, governance, infrastructure management, and operational responsibility can shift back onto the organisation running the model.

When it came to the biggest hidden cost of AI infrastructure, however, there was much greater agreement, at least among the Think Tank panellists: engineering time spent on undifferentiated platform work rather than product development.

As one panellist observed, almost all of this is still new – which is part of the reason so many people are attending conferences like this one, to hear how others are approaching these challenges. The technology is changing rapidly, established patterns are still emerging, and the need for learning remains immense.

But this at least may become a solvable problem.

Knowledge gradually accretes within organisations, skills improve, and platforms mature. Over time, engineers should spend less effort solving the same foundational problems, while platforms become better able to accommodate changing capabilities without fundamental redesign.

That suggests the biggest costs of AI may eventually shift someplace else, and one likely candidate is data movement and management. AI systems have voracious appetites for data, while also generating enormous quantities of it through their outputs.

The infrastructure challenge of AI may therefore become less about simply providing enough compute, and more about understanding which costs genuinely contribute to better outcomes.

Which brings the discussion back to where it started.

Before organisations can decide whether AI is delivering value, they first need to answer the deceptively difficult question – what does good actually look like?

Leave a Reply