<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LLM on Chengyu Wang</title><link>https://chengyu.eu/tags/llm/</link><description>Recent content in LLM on Chengyu Wang</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>© 2026 Chengyu</copyright><lastBuildDate>Sun, 08 Feb 2026 03:00:07 +0000</lastBuildDate><atom:link href="https://chengyu.eu/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>Deleting My Fine-Tuning Notes: From 'Taming a Model' to 'Steering a Process'</title><link>https://chengyu.eu/posts/deleting-fine-tuning-notes/</link><pubDate>Sun, 08 Feb 2026 03:00:07 +0000</pubDate><guid>https://chengyu.eu/posts/deleting-fine-tuning-notes/</guid><description>Clearing out a year-old Obsidian folder on model fine-tuning made the shift obvious: from chasing model capability to chasing business results through workflow and RAG.</description></item><item><title>A Weekend with RAGFlow: Building My Own RAG-Based Knowledge Base</title><link>https://chengyu.eu/posts/ragflow-personal-rag-knowledge-base/</link><pubDate>Sun, 25 Aug 2024 16:00:30 +0000</pubDate><guid>https://chengyu.eu/posts/ragflow-personal-rag-knowledge-base/</guid><description>Deploying the open-source RAGFlow project, testing it against real documents, and wiring it up to local Ollama models to keep costs at zero.</description></item><item><title>Comparing Open-Source LLMs for Customer-Service Conversation Summaries</title><link>https://chengyu.eu/posts/comparing-llms-for-customer-service-summaries/</link><pubDate>Sat, 03 Aug 2024 12:00:04 +0000</pubDate><guid>https://chengyu.eu/posts/comparing-llms-for-customer-service-summaries/</guid><description>Running ten locally-hosted models on consumer GPU hardware to see which ones best summarize customer-service calls after speech-to-text.</description></item><item><title>LLM Agents 101</title><link>https://chengyu.eu/posts/llm-agents-basics/</link><pubDate>Tue, 30 Jul 2024 00:00:09 +0000</pubDate><guid>https://chengyu.eu/posts/llm-agents-basics/</guid><description>What an &amp;lsquo;agent&amp;rsquo; actually means in the context of large language models, and a worked example in a customer-service setting.</description></item><item><title>Fixing a CUDA Error When Running Gemma2 in Ollama</title><link>https://chengyu.eu/posts/ollama-gemma2-cuda-error/</link><pubDate>Sat, 27 Jul 2024 22:36:33 +0000</pubDate><guid>https://chengyu.eu/posts/ollama-gemma2-cuda-error/</guid><description>A one-line fix for a CUBLAS_STATUS_NOT_INITIALIZED crash: just update Ollama.</description></item><item><title>Fine-Tuning Your Own Model with Free GPU Compute</title><link>https://chengyu.eu/posts/free-gpu-finetuning-colab/</link><pubDate>Fri, 10 May 2024 17:00:18 +0000</pubDate><guid>https://chengyu.eu/posts/free-gpu-finetuning-colab/</guid><description>After AMD&amp;rsquo;s ROCm ecosystem let me down for local fine-tuning, free Colab GPUs turned out to be the pragmatic way to fine-tune Llama 3 for free.</description></item><item><title>Running Llama 3 Locally with Open WebUI</title><link>https://chengyu.eu/posts/local-llama3-open-webui/</link><pubDate>Fri, 10 May 2024 09:00:11 +0000</pubDate><guid>https://chengyu.eu/posts/local-llama3-open-webui/</guid><description>Pairing a consumer GPU running local Llama 3 with Open WebUI and One API to combine local and remote models behind one interface.</description></item><item><title>Building a Free, OpenAI-Compatible API on Top of Cloudflare Workers AI</title><link>https://chengyu.eu/posts/free-qwen-openai-compatible-api/</link><pubDate>Wed, 01 May 2024 06:48:10 +0000</pubDate><guid>https://chengyu.eu/posts/free-qwen-openai-compatible-api/</guid><description>Wrapping Cloudflare&amp;rsquo;s free Workers AI models (including Qwen) behind an OpenAI-compatible endpoint, so existing front-ends don&amp;rsquo;t need to change a line of code.</description></item><item><title>International and Chinese Cloud Providers' LLM Landscape</title><link>https://chengyu.eu/posts/cloud-llm-landscape-2024/</link><pubDate>Sat, 27 Apr 2024 18:00:00 +0000</pubDate><guid>https://chengyu.eu/posts/cloud-llm-landscape-2024/</guid><description>A quick note on a comparison table of large-model offerings across major cloud providers.</description></item><item><title>Trying Out Local LLM Deployment: GPT-3.5 vs. Qwen-7B</title><link>https://chengyu.eu/posts/local-llm-deployment-experience/</link><pubDate>Tue, 23 Apr 2024 15:30:00 +0000</pubDate><guid>https://chengyu.eu/posts/local-llm-deployment-experience/</guid><description>Some hands-on experiments with un-tuned open models, and a few thoughts on where the real value in LLM products actually sits.</description></item></channel></rss>