Tag: AI

On-device or on-premises artificial intelligence–what does it offer

A significant direction for artificial intelligence, especially advanced AI like generative AI, is to have the AI processes performed on the same device or within the same premises as the users who will benefit from it.

This can be considered as part of edge computing because it involves the pre-processing of data before it is sent to a cloud-driven AI platform or post-processing of data coming back from a cloud-driven AI platform.

What is desireable about this is energy efficiency for cloud-based AI, reduced data transfer requirements or assurance of user privacy, corporate confidentiality and data sovereignty due to the minimum amount of data processed in an online environment. Apple even takes this further by running a private cloud specific to each Apple platform user to cater for more intense processing that can’t be performed on the device itself.

On-device and on-premises AI relies primarily on a smaller language model compared to the large language models that cloud-based AI services like ChatGPT rely on. Here they are focused on the data that exists on or is likely to come in to the machine or the logical network. Here, this cam allow for improved data management or permit a custom language model that represents personal or corporate desires.

Why on-device and on-premises AI

Samsung Galaxy AI press image courtesy of Samsung

Samsung Galaxy AI representing on-device artificial intelligence on Android mobile devices

A key desire is data security and end-user privacy. Here the data never leaves the device or premises for artificial-intelligence / machine-learning processing. This satisfies business and industry compliance expectations like privacy, corporate confidentiality and data sovereignty requirements.

Another benefit is improved performance and personalisation when it comes to artificial-intelligence processing and machine learning. Here, the processing takes place on a local machine thus avoiding the use of oversubscribed cloud computing services that can underperform under load. The AI language model ends up being highly personalised thus becoming lean.

Apple iPhone 16 press image courtesy of Apple

Apple iPhone 16 Series – first iOS device with on-device AI processing

There is reduced energy consumption compared to sending the data out to cloud-computing data centres. You also see efficient use of telecommunications links due to smaller amounts of data being sent using them as well as seeing efficient use of cloud-computing services. Another key benefit to see is improved service resilience because you aren’t heavily dependent on online resources – what if the network link fails.

Hybrid (cloud+on-device / on-premises) AI setups can allow for more sophisticated artificial-intelligence / machine-learning processing and working with multiple custom environments. This is due to having a lot of the data handling done locally before it is submitted or after results are received.

On-device AI processing

Lenovo Yoga Slim 7X 2-in-1 laptop with Qualcomm Snapdragon X Elite silicon press image courtesy of Lenovo

Lenovo Yoga Slim 7X 2-in-1 laptop with Qualcomm Snapdragon X Elite silicon – implementing on-device AI under Windows 11 for ARM microarchitecture

On-device AI is about having the AI data handled by the same device that is to make use of the data. This is facilitated through either a third processor called a neural processing unit (NPU) or a very powerful general processor that sets aside processor cores for neural processing to answer AI tasks.

A good analogy to think of are some NAS units that have a graphics processor in addition to their primary CPU. Here, these devices use the graphics processor for accelerated datatype translation like converting multimedia files in to other formats. or similar processing tasks.

There will also be an expectation to have a lot of RAM and storage capacity on these devices. This is something that is being answered easily thanks to Moore’s Law where cost of increased storage and RAM is being reduced significantly.

Such setups can be facilitated either on regular computers or mobile computing devices like smartphones and tablets.

On-premises AI processing

New Dell XPS 13 with Intel Lunar Lake Core Ultra processor press image courtesy of Dell

Dell has offered an XPS 13 laptop with Intel Lunar Lake Core Ultra CPU which has on-device AI processing for Windows 11 under IA microarchitecture

The on-premises AI approach would rely on a server or NAS on the same logical network as the end-users to process the AI data. This would come in to its own with on-premises or hybrid cloud computing setups where the desire is to keep the important data on the user’s premises.

This could represent a server or NAS that uses artificial intelligence and machine learning to make sense of a data set stored therein; or a server or NAS could perform AI tasks for client computers that don’t have on-device AI abilities. This can even lead to the creation of local chatbots that supply answers based on locally-held organisational data.

The trends associated with on-device and on-premises AI

Apple Intelligence writing tools on MacOS screenshot courtesy of Apple

Even MacOS is now supporting on-device artificial intelligence on the latest Macs with Apple silicon.

2024 has effectively become the year of general-purpose on-device AI processing with both the mobile-platform devices that run mobile operating systems and the regular computers that run desktop operating systems.

Some of the premium Android smartphones, tablets and smartwatches from the likes of Samsung and Google that are introduced in 2024 are being equipped with AI functionality. These implement Qualcomm Snapdragon mobile ARM64 processors and use this technology for voice-to-text, advanced search, machine translation, photo editing and similar functionality. Apple is introducing this kind of on-device AI to their latest iPhones and iPads powered with their latest silicon as part of Apple Intelligence, their branding of on-device AI. Here, this offers AI-driven inbox management, document and recording summarisation, image editing, AI-driven emojis and similar functions.

As well, during this year, Microsoft built in to Windows 11 on-device AI functionality which comes alive on computers that have neural-processing units. This is marketed as CoPilot+ and is being offered on laptop computers that use Qualcomm Snapdragon X (ARM64) silicon or, shortly, Intel Lunar Lake Core Ultra and AMD Strix Point (IA-64) silicon. These offer video transcription and captioning, image creation and editing, video editing, document summarisation amongst other things. This has been underscored by a  deluge of CoPilot+ AI-capable laptops being launched or given their first outing at the Internationaler Funkaustellung 2024 in Berlin with some of the units equipped with Intel silicon and others with Qualcomm Snapdragon X silicon.

Apple is also offering a similar kind of artificial intelligence for the latest Macintosh computers with the latest Apple M-series silicon. This will offer the same kind of features as their iOS and iPadOS implementations but with a richer interface. For all the Apple operating systems, there is support for hybrid ChatGPT operation with a “private cloud” arrangement to protect users’ data.

QNAP and Synology are working on equipping newer NAS units and newer versions of the NAS operating systems for artificial intelligence with AI being seen as part of a NAS’s feature set. But this will primarily be about managing or indexing data held on these devices themselves but someone even prototyped a NAS-based local ChatGPT setup as a proof of concept about on-premises generative AI setups which would then be about secure AI operations.. There will be the idea of using business or enthusiast grade NAS units as part of edge-computing setups to permit pre-processing of data before submitting to cloud-based AI.

Conclusion

On-device and on-premises artificial intelligence including hybrid setups such as edge-based AI or private cloud AI is expected to be a key turning point for this technology. This will most likely be due to a call for secure private and bespoke data handling requirements coming about and to keep generative AI technology relevant for most users.

Questions are being raised about generative artificial intelligence

What is AI

Artificial intelligence is about use of machine learning and algorithms to analyse data in order to make decisions on that data. It is more so about recognising and identifying patterns in the data presented to the algorithm based on what it has been taught.

This is primarily used with speech-to-text, machine translation, content recommendation engines and similar use cases. As well, it is being used to recognise objects in a range of fields like medicine, photography, content management, defence, and security.

You may find that your phone’s camera uses this as a means of improving photo quality or that Google Photos uses this for facial recognition as part of indexing your photos. Or Netflix and other online video services use this to build up a “recommended viewing” list based on what you previously watched. As well, the likes of Amazon Alexa, Apple Siri or Google Assistant use this technology to understand what you say and create a conversation.

What is generative AI

Generative artificial intelligence applies artificial intelligence including machine learning towards creating content. Here, it is about use of machine learning, typically from different data collections, and one or more algorithms to create this content. It is best described as programmatically synthesising material from other material sources.

This is underscored by ChatGPT and similar chatbots that use conversational responses to create textual, audio or visual material.  This is seen as a killer app for generative AI. But using a “voice typeface” or “voice font” that represents a particular person’s voice for text-to-speech applications could be a similar application.

Sometimes generative AI is used as a means to parse statistical information in to an easy-to-understand form. For example, it could be about an image collection of particular cities that is shaped by data that has geographic relevance.

The issues that are being raised

Plagiarism

Here, one could use a chatbot to create what apparently looks like new original work with material from other sources without attributing the content creators for the material that existed in these sources.

Nor does it require the end-user to make a critical judgement call about the sources or the content created or allow the user to apply their own personality to the content.

This affects academia, journalism, research, creative industries and other use cases. For example, education institutions are seeing this as something that impacts on how students are assessed, such as whether the classic written-preferred approach is to be maintained as the preferred approach or to be interleaved with interview-style oral assessment methods.

Provenance and attribution

It can also extend to identifying whether a piece of work was created by a human or by generative artificial intelligence and identifying and attributing the original content used in the work. It also encompasses the privacy of individuals that appear in work like photos or videos; or where personal material from one’s own image collection is being properly used.

This would be about, for example, having us “watermark” content we create in or export to the digital domain and having to identify how much AI was used in the process of creating the content.

Creation of convincing disinformation content

We are becoming more aware of disinformation and its effect on social, political and economic stability. It is something we have become sensitised to since 2016 with the Brexit referendum and Donald Trump’s election victory in the USA.

Here, generative artificial intelligence could be used to create “deepfake” image, audio and video content. An example of this being a recent image of an explosion at the Pentagon, that was sent around the Social Web and had rattled Wall Street.

These algorithms could be used to create the vocal equivalent of a typeface based on audio recordings of a particular speaker. Here, this vocal “typeface” equivalent could then be used with text-to-speech to make it as though the speaker said something in particular. This can be used as a way to make it as though a politician had contradicted himself on a sensitive issue or given authority for something critical to occur.

Or a combination of images or videos are used to create another image or video that depicts an event that never happened. This can involve the use of stock imagery or B-roll video mixed in with other material.

Displacement of jobs in knowledge and creative industries

Another key issue regarding generative artificial intelligence is what kind of jobs this technology will impact.

There is a strong risk that a significant number of jobs in the knowledge and creative industries could be lost thanks to generative AI. This is because the algorithms could be used to turn out material, rather than having people create the necessary work.

But there will be a want in some creative fields to preserve the human touch when it comes to creating a work. Such work is often viewed as “training work” for artificial-intelligence and machine-learning algorithms.

It may also be found that some processes involved in the creation of a work could be expedited using this technology while there is room to allow for the human touch. Often this comes about during editing processes like cleaning-up and balancing audio tracks or adjusting colour, brightness or contrast in image and video material with such processes working as an “assistant”. It can also be about accurately translating content between languages, whether as part of content discovery or as part of localisation.

There could be the ability for people in the knowledge and creative industries to differentiate work between so-called “cookie-cutter” output and artistic output created by humans. This would also include the ability to identify the artistic calibre that went in to that work.

The want to slow down and regulate AI

There is a want, even withing established “Big Tech” circles, to slow down and regulate artificial intelligence, especially generative AI.

This encompasses slowing down the pace of AI technology development, especially generative AI development. It is to allow for the possible impact that AI could have on society to be critically assessed and, perhaps, install “guardrails” around its implementation.

It also encompasses an “arms race” between generative-AI algorithms and algorithms that detect or identify the use of generative AI in the creation of work. It will also include how to identify source material, or the role generative AI had in the work’s creation. This is because generative AI may have a particular beneficial role in the creation of a piece of work such as to expedite routine tasks.

There is also the emphasis on what kind of source material the generative AI algorithms are being fed with to generate particular content. It is to remind ourselves of the GIGO (garbage in, garbage out) concept that has been associated with computer programming where you can’t make a silk purse out of a sow’s ear.

What can be done

There has to be more effort towards improving social trustworthiness of generative AI when it comes to content creation. It could be about where generative AI is appropriate to use in the creative workflow and where it is not. This includes making it feasible for us to know whether the content is created by artificial intelligence and the attribution of any source content being used.

Similarly, there could be a strong code of ethics for handling AI-generated content especially where it is used in journalism or academia. This is more so where a significant part of the workload involved in creating the work is contributed by generative AI rather than it being used as part of the editing or finishing process.