Cloud Services
Public, open-weight or dedicated? A CIO’s guide to the next phase of university AI
-
Karl Napper
- 5 min read
Summary
University AI should not be a choice between public and private infrastructure. Start with public frontier models to experiment, then optimise each production workload based on capability, cost, control, sovereignty and demand. Open-weight models can provide greater flexibility without requiring universities to own GPUs, while dedicated capacity only makes sense when workloads are predictable, sensitive or large enough to justify it.
The goal is simple: use the right model, with the right level of control, on the right infrastructure and economics for each workload.
Public AI is the natural place to start. But as experimentation turns into production, university CIOs face a different challenge: choosing the right model, level of control and infrastructure for each workload.
For most universities, the first phase of generative AI has been about access.
Copilot, ChatGPT, Claude, Gemini and public AI APIs make it easy to experiment without forecasting GPU demand or operating AI infrastructure. That remains the right starting point.
The question changes when an experiment becomes a workload. An assistant starts answering thousands of requests. AI becomes embedded in admissions, finance or IT operations. An agent begins interacting with institutional systems. What started as occasional consumption becomes something measurable.
Does every production workload still need the model and consumption model we used to prove it?
Start on the frontier, then measure
Using highly capable frontier models during experimentation makes sense. They remove capability as a constraint while teams establish whether a use case actually works.
Once the workload is proven, the questions change. How often does it run? What type of reasoning does it require? What data does it process? What does a completed task cost? Is demand predictable? Does it need the strongest available model?
Gartner’s Australian research points to organisations moving from experimentation towards production inference, with greater use of smaller domain-specific models and hybrid architectures.
Start with the capability needed to prove the idea. Then right-size the model for the workload you actually built.
A complex reasoning task may genuinely need a frontier model. A service repeatedly classifying documents, extracting fields, retrieving policy or executing well-defined tools may not.
The question is not whether an open-weight model is generally as capable as the frontier. It is whether it reliably meets the quality threshold for that specific workload.
Open-weight does not mean buying GPUs
This is where open-weight models become useful. Their value is not simply that they are ‘open’. They provide greater choice over where inference runs, how models are operated, model lifecycle and the controls surrounding the workload.
But there is a large gap between calling a public API and operating your own GPU infrastructure.
- Public or shared consumption-based inference: Best suited to experimentation, variable demand or workloads where advanced frontier reasoning remains valuable.
- Dedicated managed open-weight inference: Consider when a repeatable production workload needs greater isolation, control, sovereignty or predictable capacity.
- Dedicated GPU capacity: Consider when there is a sustained portfolio of AI workloads that genuinely requires infrastructure-level control and can make productive use of the capacity.
These are not maturity levels. A university may use all three at the same time.
MCS has structured Launch® AI around the same principle. Shared Model-as-a-Service provides a low-commitment path for variable demand. Dedicated Model-as-a-Service provides private, isolated capacity for production workloads requiring greater control. Dedicated GPU and container services provide deeper infrastructure control where the workload genuinely warrants it.
The other important capability is model routing. Rather than asking one model to do everything, requests can be routed according to capability, cost, modality or sovereignty requirements. Frontier models can remain part of the architecture without becoming the default for every request.
Sovereignty and economics should follow the workload
As AI moves deeper into university operations, sovereignty becomes increasingly relevant. An employee summarising public content has very different requirements from an application processing student information, employee records, security data or other sensitive institutional information.
The questions should go beyond where the GPU sits. Where does inference occur? Where do prompts and retrieved data travel? Who operates the environment? What is retained? Which jurisdictions are involved? Can the underlying model be changed without redesigning the application?
For many workloads, public AI will remain entirely appropriate. Others may justify Australian-hosted open-weight inference. A smaller number may require dedicated capacity. Sovereignty should be one factor in workload placement, alongside capability, performance, cost, integration and operational importance.
The same applies to economics. Consumption pricing works well when demand is uncertain. Dedicated capacity becomes worth considering once utilisation is predictable. The better question is not simply ‘what is our token price?’ It is: what does it cost to complete the task?
That might be cost per application processed, enquiry resolved or workflow completed. Once that is measurable, CIOs can sensibly compare public consumption, shared inference and dedicated capacity.
Universities have one additional capacity question
Universities also have something most enterprises do not: significant research computing demand.
Research AI and institutional AI have different funding, governance and workload characteristics and should not simply be combined. But as both create greater GPU demand, it is reasonable to ask whether the underlying capacity always needs to remain separate.
Could different security and allocation models share infrastructure? Could steady institutional inference complement variable research demand? Or do grant, governance and service requirements make separation preferable?
MCS will be exploring that question at eResearch Australasia 2026 in Sharing the GPU: Can Research and Institutional IT Build AI Capacity Together? with researchers, university IT leaders and shared infrastructure operators.
Right model, right design, right economics
The destination is unlikely to be either public AI or private AI. It will be a combination.
Use frontier models where their additional intelligence creates value. Use open-weight models where they meet the workload requirement with greater control or better economics. Reserve dedicated capacity when utilisation, isolation or sovereignty justify it. Invest in dedicated GPUs when there is enough sustained demand to support them.
The model should be a component of the architecture, not the architecture itself.
The next phase of university AI is not about choosing one platform. It is about having enough choice to put the right model, with the right level of control, on the right economics for each workload.
Five questions for CIOs as an AI workload matures
- Does this workload genuinely require a frontier model?
- Is demand predictable enough to measure and optimise?
- Do its data or sovereignty requirements justify greater control?
- Would shared managed inference meet those requirements, or is dedicated capacity genuinely necessary?
- Is there enough sustained demand across the institution to justify dedicated GPUs?
Karl Napper
Karl Napper
How can we help?
From dedicated infrastructure to managed private cloud specifically designed for Australian higher education and research
Launch® AI
Sovereign AI platform
CAUDIT Cloud
Cloud infrastructure for higher education
Managed Private Cloud
Dedicated infrastructure, workload placement, sovereignty
Contact
Find the Right Home for Your AI Workloads
Talk to our AI specialists about the right mix of public AI, open-weight models and dedicated infrastructure for your university’s workloads.
- 1800 004 943
- 15/2 Market Street, Sydney NSW 2000 AU







