Thursday, 23 January 2025
25.4 C
Singapore
19.5 C
Thailand
21.4 C
Indonesia
26 C
Philippines

The founder says Chinese AI can thrive with bigger models and more data

Stepfun's founder champions scaling laws and multimodality in AI development, predicting a trillion-parameter model revolution in China's AI industry.

If you follow the latest developments in artificial intelligence, youโ€™ll find that bigger models and more data are the keys to success. Jiang Daxin, the founder of Stepfun, a Shanghai-based AI start-up, believes in the power of scaling laws in large language model (LLM) development. Despite challenges like lower investment and a lack of advanced chips in China, Jiang remains optimistic.

Jiang, who used to work at Microsoft, shared his thoughts at the World Artificial Intelligence Conference (WAIC) in Shanghai. He predicts that LLMs will eventually reach hundreds of trillions of parameters, greatly enhancing their capabilities.

The promise of scaling laws

Scaling laws are all about the relationship between an AI modelโ€™s performance and its number of parameters. Generally, larger models perform better, especially with more data and excellent computational resources, although the improvements can slow down after a certain point. Big tech companies invest heavily in advanced technology, particularly Nvidiaโ€™s H100 chips, to maximize performance.

Jiang highlighted this trend in his talk. โ€œThe advancements in OpenAIโ€™s GPT series, which powers ChatGPT, and the massive investments in supercomputing centers by companies like Amazon, Microsoft, and Meta show that scaling laws work,โ€ he said on Saturday. However, he cautioned that the availability of data, skilled personnel, and concerns about return on investment could affect the pace of these advancements.

Since OpenAI launched ChatGPT in late 2022, Chinese tech giants and start-ups have been eager to develop their LLMs. China has over 200 AI models, including Alibabaโ€™s Tongyi Qianwen and Baiduโ€™s Ernie. Alibaba owns the South China Morning Post, which reported this news. Yet, many Chinese AI firms struggle to match the spending power of their US counterparts and focus instead on revenue-generating applications.

Stepfunโ€™s innovative models

Founded in April 2023, Stepfun has been dedicated to developing fundamental models. At WAIC, the company launched Step-2, a trillion-parameter LLM, along with the Step-1.5V multimodal model and the Step-1X image generation model.

Jiang also emphasized the importance of multimodality in creating a comprehensive AI. Multimodal models can process visual and other data types to develop internal representations of the external world. He explained that Stepfun aims to combine generative and comprehension abilities in a single model.

Stepfun also offers consumer-facing products, such as Yuewen, a ChatGPT-like personal assistant, and Maopaoya, an AI companion that can take on various character personalities.

The future of AI investment

โ€œLast year, global AI investments reached US$22.4 billion, with 70 to 80 percent going to companies developing large models,โ€ said Alex Zhou Zhifeng, managing partner at Qiming Venture Partners, at another WAIC side event. Qiming was an early investor in Stepfun.

Zhou noted that more investments in AI applications are expected soon, partly due to decreasing token costs. In AI, a token is a basic data unit processed by algorithms.

Peng Wensheng, an economist at China International Capital, added that Chinaโ€™s AI model market is projected to reach about 5.2 trillion yuan (US$715.1 billion) by 2030. The size of the size of the industrial AI market is expected to be around 9.4 trillion yuan.

This optimistic outlook suggests a bright future for AI development in China, driven by the potential of scaling laws and innovative models like those from Stepfun.

Hot this week

ASUS IoT edge AI computers enhanced with NVIDIA Jetson Orin technology for improved AI performance

ASUS IoT edge AI computers with NVIDIA Jetson Orin now support Super mode, boosting generative AI performance by up to 2X with the latest JetPack SDK.

DeepSeek claims its ‘reasoning model’ outperforms OpenAIโ€™s o1 on key benchmarks

DeepSeekโ€™s R1 claims to outperform OpenAIโ€™s o1 in reasoning tasks, but regulatory and geopolitical issues shape its limitations and potential impact.

Perplexity AI proposes merger with TikTok US

Perplexity AI submitted a merger bid for TikTok US, aiming to integrate video into its AI search engine before the ban deadline.

ASUS unveils ProArt PA401 Wood Edition PC case

ASUS launches the ProArt PA401 Wood Edition PC case with superior cooling, sustainable ash wood design, and user-friendly assembly features.

OPPO partners with football prodigy Lamine Yamal as global ambassador

OPPO announces Lamine Yamal as global ambassador, combining football and technology to inspire young people through the "Make Your Moment" campaign.

Garmin launches Instinct 3 Series smartwatches with AMOLED displays

Garmin unveils the Instinct 3 Series, rugged smartwatches with AMOLED displays, solar charging, advanced health monitoring, and military-grade durability.

UK unveils digital wallet and AI chatbot to revolutionise public services

The UK announces a digital wallet for IDs and an OpenAI-powered chatbot to enhance public services, aiming for secure and efficient solutions.

Apple set to launch iPhone SE 4 with Dynamic Island and iPad Air featuring M3 chip

The iPhone SE 4 with Dynamic Island and iPad Air with M3 chip are expected to launch soon. They will offer modern design and performance upgrades.

President Trump signs executive order delaying TikTok ban for 75 days

Trump delayed the TikTok ban with a 75-day executive order, allowing time to address national security concerns and find a resolution.

Related Articles