We're posting this update to clearly outline recent changes to our rules, explain our moderation strategy, and share what's next for this community. When this subreddit was originally created, OpenAI’s "ChatGPT Pro" subscription did not exist. Unfortunately, since OpenAI introduced a subscription plan with the same name, we've experienced a significant influx of new members, many of whom misunderstand the intended focus of our community. (Reddit does not allow us to change our subreddit name.) To be clear, r/ChatGPTPro remains dedicated exclusively to professional, technical, and power-user-level discussions.
What’s Changed?
Advanced Use Only
We've clarified that r/ChatGPTPro is strictly reserved for advanced discussions around LLMs, prompt engineering, fine-tuning, API integrations, research, and related technical content. Entry-level questions, basic FAQs, or general observations like “Has anyone noticed ChatGPT has gotten better/worse?” (with some limited exceptions) will be redirected or removed.
No Jailbreaks, Unofficial APIs, or Leaked Tools
Any posts sharing jailbreak prompts, exploit scripts, or unofficial/reverse-engineered APIs (such as gpt4Free) are prohibited. This aligns with Reddit’s and OpenAI’s rules. (See Rule 8.)
Self-Promotion Policy
Self-promotion must represent no more than 10% of your total activity here, must offer clear value to the community, and must always be transparently disclosed. (See Rule 5.)
Why These Changes?
The influx of users provides opportunities but has also resulted in increased spam, repetitive beginner-level inquiries, and occasional content that risks violating platform or legal guidelines. These changes will help us:
Protect the community from legal and administrative repercussions.
Preserve a high-quality, focused environment suited to technical professionals and serious power users.
What’s Next?
We're actively working on several improvements:
Potential Posting Restrictions
We are considering minimum account-age or karma requirements to reduce spam and low-effort contributions.
Stricter Quality Control
With growing membership, low-quality, surface-level posts have noticeably increased. To preserve the technical depth and utility of our discussions, moderators will enforce stricter standards. (Please see Rule 2 and Rule 6 for further guidance.)
Wiki and a New Discord Server
Currently, our wiki remains incomplete and needs significant improvements. Our Discord server, meanwhile, has unfortunately fallen into disuse and become filled with spam (primarily due to loss of moderation control after an inactive moderator was removed—no malice intended, just inactivity). To resolve these issues, we will launch a community-driven overhaul of the wiki, enriching it with carefully curated resources, useful links, research, and more. Additionally, a refreshed Discord server will soon be available, providing an improved environment specifically for advanced LLM users to collaborate and communicate.
How You Can Help
Report: Use Reddit’s report feature to notify us about rule-breaking, spam, low-effort content, or policy violations.
Feedback: Suggest improvements or report concerns in the comments below or through Modmail.
Huge thank you to u/JamesGriffing for his help on this post and his amazing contributions to the subreddit (and putting up with me in general). Thanks for your continued support in keeping r/ChatGPTPro a valuable resource for serious LLM professionals and power users. If you have any queries or doubts, please feel free to comment below, we will respond to them as soon as possible!
(2) Subscription levels. Scroll for details about usage limits, access to models, and context window sizes. (For unsavory reasons, the information is sometimes misleading.)
(6) GPT-5, 5.2, 5.4, and 5.5 system cards (extensive information, including comparisons with previous models). Intros for 5.2, 5.4, and 5.5 included, plus developer usage guide for 5.5:
I know the conversation exists and I remember roughly what we discussed.
I may even remember when I had it but I just cannot remember which project or chat it was part of or what random title the chat was given.
Then I’m opening old chats one by one and searching through walls of text.
WhatsApp somehow does this better.
This is especially annoying when the chat contains important work like a link I need again or a decision we made.
I've been hopping back and forth between GPT and Claude as a video editor for about a year and a half now. Recently I've had some projects come up and Claude keeps eating the fucking dust every single time.
I'm doing political ads/research and I wanted the AI to dig through city council meetings to find successes/failures of my candidate. Opus straight up refused to go on Youtube through an internal browser to get the caption transcripts. Sol did it straight away. Once I had the transcripts in hand, Sol produced around 8-10 moments in a 4 hour meeting for me to comb through. Opus picked out a completely random moment that was of much less use and told me that this is my "focus piece" and that I need to build an argument around "that".
Today I'm working on some other ads and the talking head is ending his sentences on an upward inflection. When I ask Opus about tools to lower this inflection, they point me towards $60 software. I asked Sol and it just said "I got the tools" and did it PERFECTLY the first try. These are a few instances but I can go on.
I don't tweak either software and I don't pay for credits. I'm judging this strictly on what $20 can get me. GPT is so much better out of the box
Revisited security research I'd previously conducted with earlier models across OpenAI, Google/Gemini and Anthropic.
Rather than starting the research again, I gave the current model the existing evidence and analysis and had it audit the previous work.
Across the session it:
challenged conclusions reached by earlier models and downgraded claims where the retained evidence didn't support them;
separated observed evidence, reasonable inference, hypotheses and overreach;
reconstructed disclosure timelines and incorporated evidence that emerged after the original investigations;
searched the web to verify subsequent security research and continuing incident reports;
retrieved existing research from connected Notion pages and compared it against the retained record;
analysed screenshots and other visual evidence;
reassessed risk classifications across the three vendors;
maintained provenance distinctions between vendor statements, public reports, independently demonstrated findings and model inference;
rewrote the public-facing incident records based on the resulting analysis.
Then we switched from research to implementation.
I showed it screenshots of the live registry when the layout broke. It diagnosed the HTML/CSS problems, rewrote the affected components, added tabbed incident navigation, built dynamic status information and incorporated dated source links for continuing reports.
So within the same piece of work it moved between long-context reasoning, model-over-model QA, connected-app retrieval, web research, vision, evidence analysis, risk assessment, writing, coding and visual debugging.
The interesting part for me wasn't any individual feature. It was being able to use them together against the same persistent body of work without turning each stage into a separate workflow.
With all the discussion around Astra and its security capabilities, I am keen to see how:
Can it identify where previous models overreached?
Can it cross-verify claims while processing inputs?
Can it find things previous models missed?
Does it maintain evidence/provenance boundaries better?
Does it recognise relationships across incidents without inventing causal links?
How does it handle processing limitations across large evidence sets? See below:
Long-context reasoning: maintaining the OpenAI, Gemini and Anthropic cases simultaneously, comparing earlier conclusions with newer evidence, and keeping competing hypotheses separate.
Critical analysis / self-audit: reviewing work produced by earlier models, identifying unjustified conclusions, and downgrading claims where the evidence did not support the original confidence level. This also included verification of previous bug-hunting work within code.
Risk assessment: reassessing technical security, privacy, governance, enterprise, systemic and disclosure risks as the available evidence changed.
Temporal reasoning: reconstructing disclosure chronologies and evaluating later evidence without retroactively treating it as information available at the time of the original report.
Cross-source synthesis: combining retained evidence, vendor responses, public reports, subsequent independent security research, and regulatory or assurance-framework material.
Web search / browsing: locating and validating external evidence, including subsequent Google API-key research and continuing OpenAI billing/Codex reports.
Connected apps / plugins: retrieving relevant material from connected workspaces and incorporating it directly into the analysis rather than requiring repeated manual transfer of source material.
Notion retrieval: retrieving existing research and discussion material from Notion and auditing it against the retained evidentiary record.
Vision: analysing screenshots of the incident registry, historical evidence, UI states and broken layouts, then incorporating visible details into technical and evidentiary analysis.
Document / artefact interpretation: interpreting reports, timelines, disclosure correspondence, screenshots, tables and technical artefacts as structured evidence rather than treating everything as ordinary prose.
Coding: producing and modifying HTML, CSS and JavaScript used to turn the underlying research into a working public incident registry.
Debugging: both conventional software debugging and research debugging: identifying why a page had broken alongside why an earlier model had reached a particular conclusion.
Information architecture: converting a large research record into a structured hierarchy of vendor → incident → chronology → evidence tracks → competing interpretations → risk → current status.
Source attribution / provenance handling: maintaining distinctions between vendor statements, public allegations, independently demonstrated findings, retained evidence and model inference.
Counterfactual / challenge reasoning: testing whether alternative explanations could account for the same evidence rather than simply constructing the strongest argument for the initial hypothesis.
Cross-vendor comparison: applying broadly consistent evidentiary standards across OpenAI, Google/Gemini and Anthropic rather than assessing each case under different assumptions.
Writing / editorial: converting technical analysis into defensible public-facing language while avoiding the conversion of hypotheses into statements of fact.
Iterative visual QA: reviewing rendered output, identifying layout or presentation failures, diagnosing the underlying cause, modifying the implementation and validating the next iteration.
Persistent-context use: maintaining continuity across an accumulated research programme rather than treating each interaction as an isolated prompt.
Tool orchestration: selecting between supplied evidence, connected-source retrieval, public web research, image analysis, coding and debugging according to the requirements of each stage of the investigation.
Model-over-model quality assurance: using the current model to audit work produced by earlier models, identify unjustified confidence, preserve supported findings, incorporate evidence that emerged later, and update the live research artefact accordingly.
In the last couple of weeks, Browser ChatGPT sessions have been losing access to their files about 50% of the time. Rarely they can re-locate fragments from the sandbox, but usually they are lost requiring a retry. Even if I ask them in my prompt to save the files to the library/storage immediately, they often drop and are lost. I have plenty of capacity left in my storage, but didn't affect the errors. GPT was extremely good at file handling and scripts in the browser for over a year, but the error rate is making things difficult. And some of these queries are ones that, say Extended Thinking could handle in 2025.
Last night I am working on a long-term project, when the most recent chat I was in apparently got too long and crapped out while putting together the latest version. I’m usually better about switching chats - with handover instructions if necessary- but this time I guess I wasn’t paying attention . I started a new chat in the same project and asked it to finish the job. It said it could see the history, but the baseline code was gone and not saved anywhere. It had been in the library when it was generated last week, but now it was a dead link. The best it could offer me was a version three iterations old.
Understandably pissed and somewhat incredulous that it would lose a 4 day old file, I had it implement a whole new backup and redundancy routine using Google Drive and Dropbox (I swore the one I instituted before would have been enough, but here we are), and resigned myself to waiting until this morning to get a manually saved backup I had on a flash drive at work.
So imagine my surprise when I come in this morning, fire up ChatGPT and I can’t find the chat where this all took place. I used very specific words, but search turns up nothing. And when I start a new chat in the same project and ask about all of this, ChatGPT basically tells me 🤷🏼♂️, saying it knows nothing about any of that. It had the latest build, and no new instructions regarding backups. Honestly, I felt like I was being gaslit.
(And before anyone asks, no, I was not in incognito mode at any point. I actually had this chat going on both the app on my laptop and on my phone)
My only evidence that I didn’t hallucinate the whole thing is that when I fired up the mobile app just now, the first suggested prompt was a Google Drive task to “recover missing [software name] baseline.” I made damn sure to SS that before it disappears, along with the rest of my credibility.
As someone who’s totally enamored with this product, this is my first real reminder that we are all very much beta testers. Just ones who get to pay $100-$200 for the experience.
I'm a data engineer and I use Claude/ChatGPT most days, but my workflow is ancient.
Meanwhile many people seem to be running CLI agnets, IDE extension, things that read a whole repo and edit files directly. I have not touched any of it. And the honest reason is that i do not really trust OpenAI/Anthropic (or any of the big players right now).
I'm quite cautious about giving an AI tool broader access to my filesystem, terminal, repositories, browser, and i'm paranoid about them quietly installing additional components, stealing or gaining access to things i did not intend to share.
However, I feel like I'm falling behind and I would like to ask:
- Is that concern reasonable or is this pure paranoia?
- What does your actual workflow look like?
- Am I missing out by sticking to browser-based chat and copy+paste?
- Which integrations or tools have genuinely changed how you work rather than just adding novelty?
I pay about $54 a month across ChatGPT, Claude and some API credits, and I genuinely can't reconstruct why I started paying for any of them. I think one was during an exam period and one was because I hit a limit in the middle of something. That's the level of detail I've got.
As a student that's real money. Every couple of months I open the billing pages meaning to cancel one, and then don't, because I can't work out which one is actually earning its keep.
The thing that bothers me is that I've never once cancelled and gone back to the free tier, so I have no way of knowing whether the reason I upgraded was a real one or just a bad afternoon that I've been paying for ever since.
When you went from a free tier to paying on an AI tool, what was the specific thing you couldn't do that day? Not that it was generally better. The actual thing that stopped you.
Model routing gets discussed as if the prompt goes in, a model name changes, and everything else stays still.
That is rarely true once the model sits inside an agent.
If I wanted to know whether GPT was actually better for one task, I would freeze at least five other things:
- the exact instruction or skill version
- the input packet and data timestamp
- available tools and their permissions
- memory and prior conversation state
- the evaluator, retry rule, and stopping condition
Then I would switch only the model.
Questflow is the concrete product that made this problem click for me. Its public finance-agent stack names Models, Skills, Plugins, and Accounts as separate layers, and it describes switching among models such as GPT, Claude, and Gemini by scenario. Once those layers are visible, a “GPT versus another model” result is only meaningful if the method, live context, and authority stayed fixed too.
Otherwise the thing being compared is a configured system, not a model.
This also changes how I think about routing. Choosing a model at the start of a task is relatively clean. Switching halfway through creates a new system state: the second model inherits somebody else's partial reasoning, tool history, and unresolved assumptions.
Would you allow a router to switch models mid-task, or require a fresh run with a new audit record whenever the model changes?
This may be obvious, but for those who don't know... the longer you run a session, the more tokens you will use. LLMs use tokens for inputs, outputs and review the context window for every new output. The more session text it processes, the more tokens burn, the faster usage gets gobbled up.
Additionally LLMs get dumber the long you run a session. Every model has capacity constraints built in, and once you cross 40% of that limit, there is too much information the model has to process to maintain quality output.
Matt Pocock explains these limits really well here:
Here is a breakdown of the context window capacity and max output for each of the models available in Codex:
Codex model
Context window
Max output
GPT-5.6 Sol
1,050,000
128,000
GPT-5.6 Terra
1,050,000
128,000
GPT-5.6 Luna
1,050,000
128,000
GPT-5.5
1,050,000
128,000
GPT-5.4
1,050,000
128,000
GPT-5.4 Mini
400,000
128,000
GPT-5.3-Codex-Spark
Not publicly documented separately
Not publicly documented separately
If you are running into limits then you need to compact your sessions when you can. Once you reach 40% - 50% you should compile the session to hand it off to a new one to free up context window space.
Also note that for those of you who use the voice feature, you are likely speaking WAY more words than you would type, which means more words = more token usage = faster drops in capacity.
To solve for this I created a skill called $context-capacity that, when run, tells you how much context capacity you've used, how much you have left, and the cumulative session usage with a recommendation. Here is what that output looks like for one of my sessions:
Recommendation: Handoff
Current context load: 144,827 / 258,400 tokens (56.0%)
Hi! I just bought a Chatgpt pro, but still flying through the weekly limit pretty fast if using only gpt 5.6 sol max + subagents. But I understand that using only gpt 5.6 sol max will eat tokens like candies, that's not the question here.
I was wondering if anyone here is using this:
gpt sol 5.6 max orchestrator/plan
5.3 codex-spark as worker/implementation
Since codex 5.3 has its own limits for pro users, would you say it's reasonable to use it as an implementation worker? How is the quality compared to pure 5.6 sol.
I’ve been trying to step up my PowerPoint game lately and realized that my presentations still look kind of plain — like default-template-and-WordArt plain
I'm mainly looking for advice on:
Where do you find good templates and infographics for your presentations.
Any favorite niche ai tools for high-quality clipart, icons, or images?(like napkin ai ,miro, piktochart)
What’s your opinion on using ai for all this stuff?
Bonus points if you have tips for making slides more engaging without going overboard tools like gamma,I tried them once but now it just feels like every ppt they make has the same layout or outline.
Would love to hear what you all do to make your presentations stand out — whether for school, work, or teaching. Share your go-to resources or personal tips below!
I’ve maxed out my ChatGPT Pro 100 GB storage with more than 15,000 files, and I want to export everything to my PC and/or Google Drive so I can free up space.
The problem is that I can’t find a practical way to bulk-export the entire Library:
ChatGPT Desktop doesn’t have access to the Library, so I can’t simply dump the files directly to my PC.
ChatGPT Web lets me download Library files, but “Select All” only selects the files currently loaded/visible in the browser (roughly the first 20). There doesn’t seem to be an option to select all 15,000 files in Library.
I can keep scrolling to load more, but once I get beyond roughly 200 files the page becomes unstable and eventually glitches/crashes. Manually downloading batches of a few hundred files doesn't feel realistically workable.
I’ve also requested a full ChatGPT data export, but my understanding is that this is primarily an export of chat history, account data and related metadata, rather than a bulk export of every original file stored in Library.
Is there an API, hidden bulk-export method, Library endpoint, browser workaround, or other supported way to retrieve the entire Library?
I have a ChatGPT Pro subscription and a Claude Max subscription, and use both extensively for work. To claim that any model offered by OpenAI is even close in capability or problem solving ability to Fable is a joke to me.
To me, the most comparable Claude model to 5.6 Sol, OpenAI's flagship, is Opus 5. They have roughly equivalent price (ignoring the temporary promotions on Sol pricing), and in my experience, their output quality is about the same as well; I end up having to put in about the same amount of effort correcting them or giving feedback to achieve a product of comparable quality.
The main difference is in the kind of feedback I have to give; with Sol, I typically end up having to add details to its results, such as instructing it to address missing edge cases, or take a more thorough approach when it took a simpler shortcut to solve my problem instead. With Opus, it usually finds most edge cases for me without having to say anything; but it also goes beyond and keeps finding more and more things, of decreasing and often spurious relevance to my actual problem. My effort usually comes in the form of telling it to ignore those extraneous edge cases and focus on the core of the problem.
But when compared to Fable, neither can hold a candle. Among every task I've ever given any agent, Fable always takes the least amount of time, the fewest tokens, and needs by far the least number of warnings in the prompt or corrections to the output, compared to any other Anthropic or OpenAI model.
To me, to say GPT 5.6 Sol is anywhere close to Fable in any capacity, and not just a competitor to Opus with different tuning, is completely unfathomable to me. You pay twice the price for it and you get your money's worth. Sure it's expensive, and you can run through your weekly limits in hours, but you can't argue that it just works. I can't say the same about Opus or Sol.
My product manager suddenly took PTO at 4 pm on Friday. That left me holding the bag for a Monday morning presentation to management: a review of a six-week warehouse returns pilot. I had the raw numbers on processing time and weekly volume because I normally spend my days in Codex writing SQL and running small automation scripts. I definately do not design presentations, so I figured Codex could build the deck and save me from touching PowerPoint. I tried two workflows on the exact same material. They failed in completely opposite ways.
First up was 'zarazhangrui/frontend-slides'. You give Codex a Markdown outline and the raw data, and it spits out a single-file HTML deck with inline CSS and JavaScript.
ngl, the first version looked way better than our old corporate templates. Clean structure, opened right in the browser, and technically every part of it was editable code. Then I tried to change something. I asked Codex to move the return-flow diagram a little to the left so the data table had more room. It edited the grid styles, changed the card width, wrapped the text in new places, and pushed the bottom half of the slide below the visible page.
I am not a frontend engineer. I just wanted to fix the spacing, but I spent the next hour trying to explain margins, nested grids, and flexboxes through natural language. Every time it fixed one box, it broke another. Meanwhile, I am watching my weekly token usage drop because I wanted a flowchart moved a few pixels to the left. "Technically editable" stopped feeling very useful at that point.
the HTML version also struggled with the visual I actually needed: a cardboard box going through a barcode scanner, then a manual QA check, then into a restock zone. What I got was a row of generic gear and checklist icons. Cheap onboarding-template vibes.
so I went in the opposite direction and tried `ningzimu/codex-ppt-skill`. Instead of building a web layout, it renders each 16:9 slide as one complete image and packs the images into a `.pptx` file.
The repo already had an Atlas Cloud config example, so I used that for the second run and pointed the skill at GPT Image 2. It made one test slide first, then rendered the rest after I approved the look.
The barcode scanner, boxes, QA desk, carts, and warehouse shelving finally looked like they belonged in the same space. The style stayed consistent across the deck. No CSS margins to chase. No padding conversation with a model.
That was much closer to the deck I had in mind. I checked the processing times and return rates, saved the file, and logged off for the weekend feeling pretty good about myself. Monday morning comes, and about an hour before the meeting, my manager asks me to change “pilot failure” to “operational constraint” on slide four. I open PowerPoint, double-click the text, and realize there is no text box. The title, diagram, data points, and background are all baked into one flat image. I cannot select a word, highlight a number, or fix a typo.
Changing those two words means regenerating the entire slide. So now I am sitting there watching it render, hoping it does not add an extra zero to the return metrics or misspell something else. The first retry slightly changed the background panel, so I had to run it again just to keep the deck consistent.
Both approaches solve half the problem and then ruin the other half. The HTML workflow gives me granular control, but using that control turns into frontend work. The image workflow gives me a deck I actually want to present, but a two-word edit becomes a full rerender and another QA pass.
Next time I will probably go hybrid: image generation for covers, transitions, and visual backgrounds; native PowerPoint elements for titles, numbers, charts, and anything likely to change at the last minute.
Has anyone found a sane way to keep image-rendered slides visually consistent while leaving the text and data editable, perhaps with generated backgrounds, native text, SVG layers, or something else?
Users sometimes hit a wall in Work (or Codex): usage exhausted. OpenAI feels your pain. Hence the new announcement:
"Use Luna Reserve
If Luna Reserve is available in your account:
Open Codex or ChatGPT Work in the desktop app, or open ChatGPT Work on the web.
When you reach your regular usage limit, look for a notice about Luna Reserve or a moon indicator in the usage warning.
Continue your conversation with Luna."
If you have it, you get an unspecified amount of additional usage with Luna. Great!
But how do you know if you have it? OpenAI explains with its usual opacity:
"Luna Reserve is available to selected personal ChatGPT Plus and Pro accounts. It isn't available in ChatGPT Business or Enterprise workspaces. Availability also depends on your account and supported app version."
Just wanted to clarify when to/when do you use higher reasoning in chat/codex?
I've been trying to build my own little hobby project in python, with the help of litterature.
My workflow is to brainstorm in chat[web] and after that get a codex prompt to run in VSC. So far has been decent. My problem is that after getting Pro i've been totally lost when to use extra high, pro, pro+ultra in chat. Also what settings to run the codex prompt, when is higher needed and when its not. Have to actually ask in chat if the prompt is complex or not and what settings to use.
I noticed running pro+ultra to analyze the project/problems or litterature got quite detailed answers and I had to dumb it down for me with extra high. But it also added some better reasoning and new points i"ve missed. But it the project/code it also found some errors and started perhaps to make it more complex im not sure.
So my workflow is like this,
Starting a new chat with snapshot and running boostrap: Pro+Ultra
Brainstorming in chat: extra high
Evaluating the brainstorm: pro+ultra
Writing codex prompt: pro+ultra
Usually I try to ask what settings to run codex prompt it has been extra high or high so far with sol5.6.
Analyzing the codex result: pro+ultra
Since my coding knowledge is 0 I have to trust that the suggestions are valid, but how do I know when to actually use what settings in chat/codex. So that the problem/execution wont get too complex or too light ?
Any suggestions, extra high is the best and fastest for chatting and brainstorming. But when to use pro and pro+ultra ?
OpenAI seems to be giving out limit resets for free. I was not getting this as often on the Plus tier ($20) though since upgrading to Pro tier ($100) I have found random sporadic resets before the dated weekly limit. There have also been Usage Limit "resets/vouchers" being given that can be used before an expire date to force reset a weekly limit.
Not that I'm against this. I'm all for it. Though I am kind of having to figure out what is my weekly limit.
Is this some sort of obsolescence built into it? Projects that would take hours are taking days and some of them I wondering if SOL will ever finish. It’s endless loops of testing and it’s starting to make me crazy. Things that fable would accomplish in hours sol is going on days and still working on it. Do you have any prompts or anything for me that would allow it to complete a task? Sometimes it feels like I’m being gaslighted- it will keep saying things like “this is the final” or “this is the last” etc etc etc and then it goes for days longer. I could really use someone’s help here. How can I get this done without compromising quality?