Optimising AI Workloads: How Bespoke Hosting Solutions Maximise Performance featured image

AI adoption has accelerated rapidly, with 88% of organisations now reporting the use of AI in at least one business function (Stanford University). However, moving from adoption to scale remains a challenge. While almost 90% of organisations are at least experimenting with AI, only 7% report having scaled it across the enterprise (McKinsey). 

Requirements can also change as an AI project develops. An environment that supports initial development and testing may not be suitable once an application moves into production, connects to live business data or begins supporting a growing number of users.

This makes understanding the workload an important starting point for AI infrastructure design. Rather than fitting AI into a predetermined platform, infrastructure can be designed around what the workload needs today, with the flexibility to adapt as those requirements change.

Different AI workloads have different infrastructure requirements

AI computing can be resource intensive, particularly for workloads involving large datasets or complex models. High-performance computing (HPC) and GPU infrastructure are therefore commonly associated with AI.

GPUs can process large volumes of data in parallel, making them well suited to many AI training and inference workloads. However, this does not mean that every AI project requires the same GPU infrastructure, or even that GPUs are always the right choice.

Depending on the model and use case, CPU-based infrastructure may also be suitable for development, testing, inference or training. Other workloads may benefit from a combination of CPU and GPU resources or an architecture spanning private cloud, public cloud and dedicated infrastructure.

The organisation deploying the AI does not necessarily need to determine these technical requirements itself. By understanding the workload, how it will be used and the requirements it needs to meet, an infrastructure provider can determine the resources and architecture needed to support it.

What changes when an AI workload moves towards production?

During development, an AI project may have a relatively limited number of users, controlled datasets and predictable resource requirements.

Moving into production can change this.

An AI application may need to connect to live data and existing business systems, support more users or meet defined expectations around performance and availability. If the application becomes part of a customer-facing service or important internal process, resilience and recovery can also become more significant.

Before moving an AI workload into production, it is therefore useful to consider:

  • how many users or applications will depend on it
  • how usage is expected to change
  • which data and systems it needs to access
  • how sensitive that data is
  • what level of performance is required
  • how important availability is to the business
  • whether there are requirements around where data is stored or processed
  • how the environment will be monitored and managed

These requirements can then be translated into decisions around compute, storage, networking, security, resilience and the wider infrastructure architecture.

Planning for changing AI requirements

It can be difficult to predict exactly how an AI application will be used once it moves into production.

An initial internal use case could expand across the organisation. A customer-facing feature could experience increasing demand. Models may change, datasets may grow and new AI capabilities may be introduced.

Infrastructure therefore needs to account not only for current requirements, but also how these could develop.

This does not necessarily mean provisioning large amounts of spare capacity from the outset. Over-provisioning infrastructure can increase costs without delivering additional value.

Instead, a bespoke approach allows resources to be aligned with current workload requirements while considering how additional compute, storage or other resources can be introduced as those requirements change.

Choosing the right environment for the workload

There is no single infrastructure model that is right for every AI workload.

Public cloud can provide rapid access to scalable resources and a wide range of AI services. Private cloud can provide greater control over the underlying environment. Dedicated HPC infrastructure can support particularly compute-intensive workloads, while hybrid architectures can allow different parts of an AI environment to run where they are best suited.

Existing infrastructure also matters. An organisation may already have applications, data or workloads running across public cloud, private cloud, dedicated infrastructure or on-premise environments.

Rather than treating AI as an isolated platform, the architecture can be designed to work alongside these existing systems.

The appropriate approach depends on the AI use case, data, performance requirements, existing environment, expected scale and the level of control the organisation requires.

Designing infrastructure around the workload

Preconfigured AI hosting can provide a straightforward route to infrastructure, but fixed configurations will not suit every workload.

For example, automatically deploying GPU infrastructure because a workload involves AI could provide significantly more compute than the project requires. On the other hand, an environment without sufficient resources could create performance issues as usage increases.

A consultative approach starts with the requirements rather than the infrastructure.

This includes understanding what the AI application needs to achieve, how it operates, the data it uses, expected demand, performance requirements and how it needs to integrate with the wider environment.

The underlying infrastructure can then be designed accordingly, with the appropriate combination of compute, storage, networking and cloud or dedicated resources.

This can help avoid unnecessary resources while providing the performance and flexibility required by the workload.

Case study: Designing bespoke infrastructure for Symbolic Mind

The importance of designing around the workload can be seen in our work with Symbolic Mind, a company developing generative AI and artificial general intelligence architecture.

Symbolic Mind needed infrastructure to support the development and testing of its EVA model. Their requirements were unique – unlike many AI models which only run on GPUs, EVA can use either CPU, GPU or both. For this development and testing project, they needed a CPU powered server. 

Other providers Symbolic Mind approached offered off-the-shelf GPU-based solutions which did not meet their requirements. Where different providers did offer a bespoke solution, the costs were prohibitive. 

They found our proposal for a bespoke 56-core CPU server was the only solution that met their needs and budget, providing the power, memory and scalability required. Following the custom build, Symbolic Mind trained their EVA LLM on the server, achieving record-breaking 3Gb per hour speeds.

The project demonstrates why the infrastructure decision should follow the workload requirements rather than assumptions about the technology being used.

Managing AI infrastructure in production

Infrastructure requirements do not stop changing once an AI application goes live.

Increasing usage, changing datasets, new models and evolving applications can all affect the resources an environment requires. Production infrastructure also needs ongoing monitoring, security management, patching and optimisation.

For organisations without the resources or desire to manage the underlying platform internally, a managed approach can move this responsibility to an infrastructure provider.

Why choose managed AI services from Hyve?

Enterprise adoption of AI at scale is still in its infancy, with new innovations requiring dynamic, flexible infrastructure.

At Hyve, our consultants work with you to understand your AI use case, existing environment and requirements before our engineers design the appropriate infrastructure. This can incorporate private cloud, HPC and other infrastructure as required, with ongoing management to help maintain performance, security and availability as workloads evolve.

Speak to one of our cloud experts today – fill out our contact form and we will be in touch.

Insights related to Blog

Do you want to know everything about Dedicated Servers?