Average Ratings 0 Ratings

Total
ease
features
design
support

No User Reviews. Be the first to provide a review:

Write a Review

Average Ratings 0 Ratings

Total
ease
features
design
support

No User Reviews. Be the first to provide a review:

Write a Review

Description

AWS Inferentia accelerators, engineered by AWS, aim to provide exceptional performance while minimizing costs for deep learning (DL) inference tasks. The initial generation of AWS Inferentia accelerators supports Amazon Elastic Compute Cloud (Amazon EC2) Inf1 instances, boasting up to 2.3 times greater throughput and a 70% reduction in cost per inference compared to similar GPU-based Amazon EC2 instances. Numerous companies, such as Airbnb, Snap, Sprinklr, Money Forward, and Amazon Alexa, have embraced Inf1 instances and experienced significant advantages in both performance and cost. Each first-generation Inferentia accelerator is equipped with 8 GB of DDR4 memory along with a substantial amount of on-chip memory. The subsequent Inferentia2 model enhances capabilities by providing 32 GB of HBM2e memory per accelerator, quadrupling the total memory and decoupling the memory bandwidth, which is ten times greater than its predecessor. This evolution in technology not only optimizes the processing power but also significantly improves the efficiency of deep learning applications across various sectors.

Description

NVIDIA DGX Cloud Serverless Inference provides a cutting-edge, serverless AI inference framework designed to expedite AI advancements through automatic scaling, efficient GPU resource management, multi-cloud adaptability, and effortless scalability. This solution enables users to reduce instances to zero during idle times, thereby optimizing resource use and lowering expenses. Importantly, there are no additional charges incurred for cold-boot startup durations, as the system is engineered to keep these times to a minimum. The service is driven by NVIDIA Cloud Functions (NVCF), which includes extensive observability capabilities, allowing users to integrate their choice of monitoring tools, such as Splunk, for detailed visibility into their AI operations. Furthermore, NVCF supports versatile deployment methods for NIM microservices, granting the ability to utilize custom containers, models, and Helm charts, thus catering to diverse deployment preferences and enhancing user flexibility. This combination of features positions NVIDIA DGX Cloud Serverless Inference as a powerful tool for organizations seeking to optimize their AI inference processes.

API Access

Has API No 

API Access

Has API Yes 

Screenshots View All

Screenshots View All

Integrations

AWS EC2 Trn3 Instances Yes 
AWS Parallel Computing Service Yes 
Amazon EC2 Inf1 Instances Yes 
Amazon EC2 Trn1 Instances Yes 
Amazon Web Services (AWS) No 
Anyscale Yes 
CoreWeave No 
Google Cloud Platform No 
Helm No 
Llama No 
Microsoft Azure No 
NVIDIA AI Foundations No 
NVIDIA Cloud Functions No 
NVIDIA DGX Cloud No 
NVIDIA NIM No 
Nebius No 
Oracle Cloud Infrastructure No 
Splunk Cloud Platform No 
WithoutBG Yes 
Yotta No 

Integrations

AWS EC2 Trn3 Instances No 
AWS Parallel Computing Service No 
Amazon EC2 Inf1 Instances No 
Amazon EC2 Trn1 Instances No 
Amazon Web Services (AWS) Yes 
Anyscale No 
CoreWeave Yes 
Google Cloud Platform Yes 
Helm Yes 
Llama Yes 
Microsoft Azure Yes 
NVIDIA AI Foundations Yes 
NVIDIA Cloud Functions Yes 
NVIDIA DGX Cloud Yes 
NVIDIA NIM Yes 
Nebius Yes 
Oracle Cloud Infrastructure Yes 
Splunk Cloud Platform Yes 
WithoutBG No 
Yotta Yes 

Pricing Details

No price information available.
Free Trial No 
Free Version No 

Pricing Details

No price information available.
Free Trial No 
Free Version No 

Deployment

Web-Based Yes 
On-Premises No 
iPhone App No 
iPad App No 
Android App No 
Windows No 
Mac No 
Linux No 
Chromebook No 

Deployment

Web-Based Yes 
On-Premises No 
iPhone App No 
iPad App No 
Android App No 
Windows No 
Mac No 
Linux No 
Chromebook No 

Customer Support

Business Hours No 
Live Rep (24/7) No 
Online Support Yes 

Customer Support

Business Hours Yes 
Live Rep (24/7) No 
Online Support Yes 

Types of Training

Training Docs Yes 
Webinars No 
Live Training (Online) No 
In Person No 

Types of Training

Training Docs Yes 
Webinars Yes 
Live Training (Online) Yes 
In Person Yes 

Vendor Details

Company Name

Amazon

Founded

2006

Country

United States

Website

aws.amazon.com/machine-learning/inferentia/

Vendor Details

Company Name

NVIDIA

Founded

1993

Country

United States

Website

developer.nvidia.com/dgx-cloud/serverless-inference

Product Features

Deep Learning

Convolutional Neural Networks No 
Document Classification No 
Image Segmentation No 
ML Algorithm Library No 
Model Training No 
Neural Network Modeling No 
Self-Learning No 
Visualization No 

Infrastructure-as-a-Service (IaaS)

Analytics / Reporting No 
Configuration Management No 
Data Migration No 
Data Security No 
Load Balancing No 
Log Access No 
Network Monitoring No 
Performance Monitoring No 
SLA Monitoring No 

Product Features

Alternatives

Alternatives

AWS Neuron Reviews

AWS Neuron

Amazon Web Services