Tag
23 articles
Deepseek plans to install 160,000 Huawei Ascend-950DT chips in an Inner Mongolia data center, marking the largest known Huawei chip cluster, though production bottlenecks may delay deployment.
OpenAI's new Jalapeño chip outperforms current industry standards in inference speed and energy efficiency, according to SemiAnalysis' InferenceX benchmark.
This article explores the complexities of AI chip performance, focusing on scaling behavior, model architecture compatibility, and how raw throughput metrics don't always reflect real-world efficiency.
AI chip startup Etched has achieved a $10.3 billion valuation despite industry skepticism, offering GPU-free solutions for AI inference.
Learn how to build and optimize AI inference pipelines using TensorFlow, mimicking the specialized approach of companies like Etched in the chip industry.
Nvidia has chosen to collaborate with its chip rival d-Matrix, combining their technologies to create a joint AI system for running inference models. The partnership marks a strategic shift toward cooperation in the competitive AI chip market.
Learn how to use ZML's open-source inference optimization software to accelerate AI model execution across multiple hardware platforms, demonstrating performance improvements through practical implementation.
Learn how to build and optimize AI inference pipelines using TensorFlow, similar to what companies like Etched are developing for specialized AI chips.
Learn to build and test AI inference systems that demonstrate how Micron's memory technology impacts AI model performance, simulating the advantages that make Micron a potential rival to Nvidia.
AI chip startup Groq is raising $650 million in internal funding as it pivots from general hardware to focus on AI inference, the process of refining how AI models respond to prompts.
London-based AI chip startup Fractile has raised $220 million to bring its in-memory computing inference chip to production, with support from Accel and Pat Gelsinger.
Meta and Stanford researchers introduce the Fast Byte Latent Transformer, reducing inference memory bandwidth by over 50% without subword tokenization.