Solo tabletop roleplaying is experiencing a golden age. Whether you are consulting fate charts in the Mythic Game Master Emulator, mapping out hexes in Cartograph: Atlas Edition, or rolling up dungeon rooms in Four Against Darkness, sometimes you need more than just dice and tables to bring the narrative to life.
By leveraging the unified memory of an Apple Silicon Mac, you can build a completely private, offline, and subscription-free AI Game Master. This guide walks you through setting up a powerful local story engine using oMLX, Meta’s Llama 3.1, and Open WebUI.
Why This Specific Stack?
To run an AI capable of acting as an immersive RPG oracle, you need three things: an efficient server, a highly capable model, and a comfortable chat interface.
- The Engine (oMLX): This is an inference server built exclusively for Apple Silicon, allowing you to manage models directly from a native macOS menu bar app. It is highly optimized and perfectly mimics the standard OpenAI API.
- The Brain (Llama 3.1 8B): We are using
Meta-Llama-3.1-8B-Instruct-4bit. In 4-bit quantization, this model only consumes about 5 to 6 GB of RAM, leaving plenty of overhead for your PDF rulebooks, virtual tabletops, and background apps. It generates text blisteringly fast and has a fantastic creative voice. - The Table (Open WebUI): This self-hosted platform provides a beautiful, ChatGPT-style web interface. It allows you to organize campaigns into separate chat threads, upload campaign notes, and even connect external tools.
Step 1: Install the oMLX Server
oMLX operates quietly in the background, serving as the bridge between your Mac’s GPU and the AI model.
- Download the macOS App directly from the jundot/omlx GitHub repository.
- Drag the
.dmgcontents into your Applications folder and launch it. - You will see a new icon in your macOS menu bar. Click it and ensure the server is actively running on port
8000. - Navigate to
http://localhost:8000/adminin your web browser to access the local dashboard.
Step 2: Download the Llama 3.1 Model
You cannot just download any model format; it must be converted specifically for Apple’s MLX framework. Fortunately, the Hugging Face community has already done this.
- In your oMLX admin dashboard, navigate to the model download or search section.
- Paste the exact repository ID:
mlx-community/Meta-Llama-3.1-8B-Instruct-4bit. - Click download. Because this is the 4-bit version, the download is only about 4.5 GB.
- Once the download finishes, click Load to push the model into your Mac’s unified memory.
GM Tip: In the oMLX settings for this model, you can create a specific profile to adjust the sampling parameters. Increasing the “Temperature” slightly (e.g.,
0.85) will make the AI’s narrative outputs more creative and unpredictable—perfect for unexpected random encounters.
Step 3: Install Open WebUI
While oMLX does have a built-in chat UI, Open WebUI is far more robust for long-term roleplaying campaigns, offering features like document uploads (Retrieval-Augmented Generation) so the AI can read your character sheets.
- Open your macOS Terminal.
- The easiest way to install Open WebUI natively is using
uv, a fast Python package manager. Run this command to install and start the server:curl -LsSf https://astral.sh/uv/install.sh | sh DATA_DIR=~/.open-webui uvx --python 3.11 open-webui@latest serve
(Alternatively, you can run it via Docker or download their standalone Desktop app). - Open your browser and go to
http://localhost:8080. - Create your local admin account (this data stays entirely on your machine).
- Click the Settings (Gear Icon) -> Connections.
- Under the OpenAI API section, add your oMLX endpoint:
http://localhost:8000/v1. You can enter any dummy text for the API key. - Click Save. Open WebUI will instantly detect the
Meta-Llama-3.1-8B-Instruct-4bitmodel you loaded earlier.
Ready to Roll
You now have a private, infinitely patient Game Master running at local hardware speeds. To kick things off, open a new chat in Open WebUI, select your Llama 3.1 model, and feed it a strict system prompt to establish the rules. For example:
“You are an expert Game Master for a solo tabletop RPG. You do not have access to the internet. Rely strictly on your internal knowledge to describe the environment, narrate NPC dialogue, and interpret oracle rolls. Ask me what my character does next at the end of every response.”