Entertainment

[2401.07013] Knowledge Distillation of Black-Box Large Language Models

[2401.07013] Knowledge Distillation of Black-Box Large Language Models

๐Ÿ”ฅ Discover this insightful post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: ๐Ÿ“Œ **What Youโ€™ll Learn**: [Submitted on 13 Jan 2024 (v1), last revised 9 Nov 2024 (this version, v2)] View a PDF of the paper titled Knowledge Distillation of Black-Box Large Language Models, by Hongzhan Chen and 5 other authors View PDF HTML (experimental) Abstract:Given the exceptional performance of proprietary large language models (LLMs) like GPT-4, recent research has increasingly focused on boosting the capabilities of smaller models through knowledge distillation (KD) from these powerful yet black-box teachers. While leveraging the high-quality outputs of these teachers is advantageous, the inaccessibility…
Read More
TOP500 at ISCโ€™26: We have a New Number 1

TOP500 at ISCโ€™26: We have a New Number 1

โœจ Discover this trending post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: โœ… **What Youโ€™ll Learn**: Hello you fine Internet folks,Here at ISC 2026 in Hamburg, Germany, we got the 67th TOP500 list where there was a surprise awaiting us. That surprise being a new Number 1 Supercomputer on the TOP500.The new number one Supercomputer on the TOP500 list is the LineShine Supercomputer in Shenzhen, China. This is the first Chinese submission to the TOP500 in 9 years and they came in swinging with a massive CPU-only system.Starting with the specifications of the CPU that powers the LineShine system, the LX2.The…
Read More
JustVugg/nanoeuler: GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT. ยท GitHub

JustVugg/nanoeuler: GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT. ยท GitHub

๐Ÿ’ฅ Check out this insightful post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: โœ… **What Youโ€™ll Learn**: A GPT-2-class language model built entirely from scratch in C/CUDA โ€” no PyTorch, no autograd, no ML libraries. The forward and backward passes are written and verified by hand, and the whole training pipeline lives in this repo: a hand-written byte-level BPE tokenizer, pretraining on a books + web corpus, and supervised fine-tuning into a chat model (RLHF/DPO planned). It runs on CPU (libm + OpenMP) for a small showcase model, and a full from-scratch CUDA engine โ€” cuBLAS matmuls, a hand-written FlashAttention, validated…
Read More
kamaludu/bash4llm: Bash-first wrapper for Groqโ€™s OpenAI-compatible API. Secure, portable, Termux-friendly. ยท GitHub

kamaludu/bash4llm: Bash-first wrapper for Groqโ€™s OpenAI-compatible API. Secure, portable, Termux-friendly. ยท GitHub

โœจ Read this insightful post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: โœ… **What Youโ€™ll Learn**: Bash4LLMโบ โ€” wrapper CLI sicuro, Bashโ€‘first e completamente auditabile per lโ€™API Chat Completions compatibile OpenAI di Groq (ed estendibile ad altri provider). Bash4LLMโบ รจ un singolo script Bash, autoโ€‘contenuto, leggibile e verificabile.Scaricalo, rendilo eseguibile, esporta la tua API key e inizia subito a usarlo. Compatibile con ambienti Unixโ€‘like: Linux, macOS, WSL, Cygwin, Termux (Android), BSD. Caratteristiche principali Lista modelli dinamicatramite GET https://api.groq.com/openai/v1/modelsโ†’ nessun modello hardcoded. Sicurezza by designโ†’ nessun uso di /tmp, nessun eval, permessi restrittivi, validazione provider avanzata. Struttura modulare a sezioniโ†’ PRECORE_BOOT, PRECORE_RUN,…
Read More

Memory Prices | DAM

๐Ÿ’ฅ Explore this trending post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: ๐Ÿ’ก **What Youโ€™ll Learn**: Historic and current memory and storage prices, collected in the spirit of John C. McCallum's classic memory-price dataset โ€” interactive, with the raw data downloadable. Hover for details, click the legend to toggle series, drag or use the slider to zoom, and use the camera icon to export an image. Price per gigabyte over time Historical lowest $/GB on a log scale โ€” one line per memory type: DRAM, NAND flash, and HBM. DRAM price by generation The DRAM line above, broken out by generation…
Read More
Authors With DRM-Free Books

Authors With DRM-Free Books

๐Ÿ”ฅ Read this must-read post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: โœ… **What Youโ€™ll Learn**: Timothy Zahn The Icarus Needle Buy on Baen Ten thousand years ago, a mysterious people known as the Icari vanished from the Spiral, leaving behind a network of portals that can instantaneously transport passengers hundreds or thousands of light-years across the stars. Gregory Roarke and his Kadolian partner Selene have been tasked with seeking out these alien artifacts and bringing them under the control of the Icarus Group. But the Groupโ€™s leadership has changed, and Roarke soon finds himself at serious odds with the new…
Read More
Working around dragons with the Lemote Yeeloong laptop and OpenBSD

Working around dragons with the Lemote Yeeloong laptop and OpenBSD

โœจ Discover this trending post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: ๐Ÿ’ก **What Youโ€™ll Learn**: Behold: the Guru of GNU! (Photo by Habib Mhenni, Wikimedia Commons, CC BY-SA 3.0.) True enlightment only comes from a truly free computing experience, probably! And while there is no nerd who lacks an opinion on Richard Stallman personally, likewise let none claim he does not practice what he preaches. Why, the very laptop in front of him was selected deliberately because it can operate with no binary blobs and no firmware you couldn't examine or replace with your own, and runs his choice of…
Read More
Does your paper really suck?

Does your paper really suck?

๐Ÿš€ Read this must-read post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: ๐Ÿ’ก **What Youโ€™ll Learn**: By A. Sina Booeshaghi ยท June 27, 2026 Oded Rechavi, at QED Science, believes that if your paper is not in the top 1% of their QED score then it "sucks". But what is this QED score and what is its purpose? Does it really measure scientific quality? If a paper is not in the 1% does it really suck? These are important questions because scientists are increasingly overwhelmed with the volume of new work posted on preprint servers and published in journals. As a…
Read More
Michigan bill would bar employers from requiring after-hours contact with workers

Michigan bill would bar employers from requiring after-hours contact with workers

๐Ÿ”ฅ Read this insightful post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: ๐Ÿ“Œ **What Youโ€™ll Learn**: A bill is pending in the Michigan Legislature that would set rules on when and for what reason an employer could contact an employee outside of a normal work schedule.ย Senate Bill 948, which was introduced by Sen. Erika Geiss, D-Taylor, has been referred to the Labor Committee. The bill is also known as the Workplace Employee Boundaries Act.ย "In an increasingly 'always-on, always available' economy, we must take action to protect workers and create stronger boundaries," Geiss saidย when introducing the bill. "Too many workers are expected…
Read More
Google limits Metaโ€™s use of its Gemini AI models, FT reports

Google limits Metaโ€™s use of its Gemini AI models, FT reports

๐Ÿš€ Read this trending post from Hacker News ๐Ÿ“– ๐Ÿ“‚ **Category**: โœ… **What Youโ€™ll Learn**: INDIA - 2025/05/13: In this photo illustration, a Meta logo is seen displayed on a smartphone with a Google logo in the background. Sopa Images | Lightrocket | Getty ImagesGoogle has put limits on Meta'sย use of its Gemini AI models after the social media company sought more computing capacity than the rival tech group could provide, the Financial Times reported on Sunday.Google, owned by Alphabet, told Meta around March it could not meet the full Gemini capacity the company had sought to purchase, the newspaper…
Read More