Gemini's free tier is getting a major downgrade from October 9 - Digital Trends
https://www.digitaltrends.com/computing/geminis-free-tier-is-getting-a-major-downgrade-on-october-9/
I guess I'll soon find out if I'm going to pay up or if 1min.ai is still working.
RC:
review my work over the last 3 months and use that to estimate workload and best usage pricing for that level
GEM:
### 1. Workload Assessment (Past 3 Months)
Based on key development streams over the last quarter, your LLM consumption breaks down into two core technical patterns:
#### A. Interactive Development & Refinement (High Output / Fast Iteration)
* **Tampermonkey Userscript Engineering:**
* Complex, stateful DOM manipulators (e.g., custom UI overhauls, stacked vertical inputs, portal/history/tray state management, URL auto-navigation).
* Intensive regex pattern matching and array sorting logic for bookmark managers.
* **3D Modeling & CAD Script Generation:**
* Parametric OpenSCAD code generation (precise geometric tolerances, rounded corner trays, custom mounting modifications).
* Orca Slicer configuration and G-code modifier troubleshooting.
#### B. Direct Clinical & Technical Reference Material (Low Output / High Precision)
* **Clinical Reference & Protocol Design:**
* Concise 25-card pocket references, micro-slip flashcards, lab panel evaluations, and patient non-pharmacological protocols.
* **System Integration & Scripting:**
* Hubitat, UPnP local streaming, and local network configuration automation.
---
### 2. Estimated Volume Breakdown
| Workload Type | Estimated Usage | Avg. Prompt Size | Avg. Output Size | Core Requirements |
| --- | --- | --- | --- | --- |
| **Userscript / CAD Code Generation** | ~150 requests / mo | ~2,500 tokens | ~1,200 tokens | Fast completion, exact code syntax without preamble |
| **Clinical Reference & Data Formatting** | ~100 requests / mo | ~1,500 tokens | ~600 tokens | Zero hallucination, strict adherence to formatting constraints |
| **Technical Troubleshooting & Regex** | ~80 requests / mo | ~1,000 tokens | ~400 tokens | Pattern matching accuracy, minimal fluff |
| **Total Estimated Monthly Volume** | **~330 requests** | **~605,000 Input Tokens** | **~272,000 Output Tokens** | **High precision, deterministic structure** |
---
### 3. API Pricing & Strategy Comparison
To replace web subscription reliance, running this volume directly through Google AI Studio / Gemini API yields substantial cost savings: emphasis mine
```
+-----------------------------------------------------------------------+
| Model Tier Pricing (per 1M Tokens) |
| - Gemini 3.1 Flash-Lite : $0.25 Input / $1.50 Output |
| - Gemini 3.8 Flash : $0.75 Input / $3.75 Output |
| - Gemini 3.1 Pro : $2.00 Input / $12.00 Output |
+-----------------------------------------------------------------------+
```
#### Cost Projection for Your Monthly Volume (~0.61M Input / ~0.27M Output)
1. **Option A: Full Gemini 3.8 Flash (Recommended Default)**
* **Input Cost:** $0.61M \times \$0.75 = \$0.46$
* **Output Cost:** $0.27M \times \$3.75 = \$1.02$
* **Total:** **~$1.48 / month**
* *Verdict:* Excellent balance for JS userscript DOM logic, OpenSCAD code generation, and complex medical/technical structuring.
2. **Option B: Full Gemini 3.1 Flash-Lite (Budget Tier)**
* **Input Cost:** $0.61M \times \$0.25 = \$0.15$
* **Output Cost:** $0.27M \times \$1.50 = \$0.41$
* **Total:** **~$0.56 / month**
* *Verdict:* Ultra-low cost for basic regex construction, simple data filtering, or quick code formatting, though less optimal for complex OpenSCAD geometry or multi-file scripts.
3. **Option C: Hybrid Model Routing (Optimal Strategy)**
* Route routine tasks (regex, formatting, basic JS helper functions) to **3.1 Flash-Lite**.
* Route heavy tasks (OpenSCAD generation, complex Tampermonkey DOM scripts, clinical reference cards) to **3.8 Flash**.
* **Total:** **~$1.00 - $1.20 / month**
---
### 4. Implementation Recommendations
1. **Leverage AI Studio API Keys:** Transition scripting workflows (e.g., local IDE extensions, desktop tools, or custom API callers) to direct API endpoints using Gemini 3.8 Flash as the primary default model.
2. **Utilize System Prompts at API Level:** Inject formatting constraints (such as direct output rules or preamble suppression) via standard API system instructions to ensure consistent execution.
3. **Context Caching:** For repetitive codebases or large reference frameworks, use context caching ($0.075 per 1M tokens) to cut input costs further during long iteration sessions.
Would you like assistance drafting a lightweight API wrapper or script configuration to automate routing across these models?
RC:
yes
GEM:
/**
* Direct Gemini API Router / Wrapper
* Target Volume: ~330 requests/mo (~0.61M Input / ~0.27M Output)
* Estimated API Cost: ~$1.00 - $1.50 / month
*/
const API_KEY = process.env.GEMINI_API_KEY;
// System instructions to enforce preamble suppression and concise execution
const DEFAULT_SYSTEM_INSTRUCTION =
"Solve using strictly verified standards. Present only executable code, final reference material, or exact matches. Omit preambles, intros, and outros.";
// Standard endpoints
const API_ENDPOINTS = {
flash_38: "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent",
flash_lite_31: "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite:generateContent"
};
/**
* Route request based on task complexity
* @param {string} prompt - Primary user query or code request
* @param {Object} options - Configuration overrides
* @param {'complex'|'routine'} options.taskType - Directs model selection
*/
async function queryGemini(prompt, { taskType = 'complex', systemInstruction = DEFAULT_SYSTEM_INSTRUCTION } = {}) {
// Option A/C Strategy: 3.8 Flash for CAD/JS/Clinical, 3.1 Flash-Lite for Regex/Formatting
const endpoint = taskType === 'complex' ? API_ENDPOINTS.flash_38 : API_ENDPOINTS.flash_lite_31;
const payload = {
system_instruction: {
parts: [{ text: systemInstruction }]
},
contents: [{
parts: [{ text: prompt }]
}],
generationConfig: {
temperature: 0.1, // Low variance for code and clinical exactness
maxOutputTokens: 2048
}
};
try {
const response = await fetch(`${endpoint}?key=${API_KEY}`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(payload)
});
if (!response.ok) {
throw new Error(`HTTP error! Status: ${response.status}`);
}
const data = await response.json();
return data.candidates[0].content.parts[0].text;
} catch (err) {
console.error("Gemini API Dispatch Error:", err);
throw err;
}
}
// Example Execution Calls:
// 1. Heavy DOM script or OpenSCAD math:
// const code = await queryGemini("Generate OpenSCAD rounded tray module...", { taskType: 'complex' });
// 2. Routine Regex or Line Filter:
// const regex = await queryGemini("Write Notepad++ regex to match date pattern...", { taskType: 'routine' });
GEM:
Would you like me to build a Tampermonkey script header wrapper to handle this API routing automatically inside your web extensions?
RC:
yes
GEM:
// ==UserScript==
// @name Gemini API Direct Router Engine
// @namespace http://tampermonkey.net/
// @version 1.0.0
// @description Direct API wrapper for Tampermonkey scripts with task-based model routing
// @match *://*/*
// @grant GM_xmlhttpRequest
// @grant GM_getValue
// @grant GM_setValue
// @connect generativelanguage.googleapis.com
// ==/UserScript==
(function() {
'use strict';
const CONFIG = {
get apiKey() {
let key = GM_getValue('GEMINI_API_KEY', '');
if (!key) {
key = prompt('Enter Gemini API Key:');
if (key) GM_setValue('GEMINI_API_KEY', key);
}
return key;
},
endpoints: {
complex: 'https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent',
routine: 'https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-lite:generateContent'
},
defaultSystemInstruction: 'Solve using strictly verified standards. Present only executable code, final reference material, or exact matches. Omit preambles, intros, and outros.'
};
/**
* Dispatch prompt to Gemini API
* @param {string} prompt
* @param {Object} options
* @param {'complex'|'routine'} [options.taskType='complex']
* @param {string} [options.systemInstruction]
* @returns {Promise<string>}
*/
window.queryGeminiAPI = function(prompt, { taskType = 'complex', systemInstruction = CONFIG.defaultSystemInstruction } = {}) {
return new Promise((resolve, reject) => {
const key = CONFIG.apiKey;
if (!key) {
reject(new Error('Missing Gemini API Key.'));
return;
}
const endpoint = CONFIG.endpoints[taskType] || CONFIG.endpoints.complex;
const url = `${endpoint}?key=${key}`;
const payload = {
system_instruction: {
parts: [{ text: systemInstruction }]
},
contents: [{
parts: [{ text: prompt }]
}],
generationConfig: {
temperature: 0.1,
maxOutputTokens: 2048
}
};
GM_xmlhttpRequest({
method: 'POST',
url: url,
headers: {
'Content-Type': 'application/json'
},
data: JSON.stringify(payload),
onload: function(response) {
if (response.status >= 200 && response.status < 300) {
try {
const resData = JSON.parse(response.responseText);
const text = resData.candidates[0].content.parts[0].text;
resolve(text);
} catch (e) {
reject(new Error('Failed to parse API response: ' + e.message));
}
} else {
reject(new Error(`API Request failed with status ${response.status}: ${response.responseText}`));
}
},
onerror: function(err) {
reject(err);
}
});
});
};
})();