Because they don't teach you this in school…

Not going to recap here, but you can certainly read my previous two posts to catch up… also if you are having trouble falling asleep.

I did promise an update on my open Autodesk ticket. I think I didn’t reply in time to something, so that ticket has disappeared.

I let CC sit for a bit after the python script purging (see last post), and things were looking pretty empty. While it was updating in the background, I had time to prep the library.

Our Revit library, like most, is divvied up by Revit version year. Like many in the past, and not so much now, each file had the year as a suffix. So Door_24.rfa and Door_25.rfa were probably the same family, just different years. One of our theories is that this duplication is what slowed down our original CC library. CC upgrades for you. So when I uploaded Door_24.rfa, it automatically created the 2025, 2026, and 2027 versions but with the name Door_24. Then, Door_25 gets uploaded and CC creates the 2026, and 2027 versions. You can see how this could spiral, and we ended up with multiple copies of the same family file.

Renaming had to include getting rid of the year suffix. Additionally, the firm changed names and abbreviations from MA to MOS. This was pretty inconsistent, but a lot of the custom families had MA somewhere in the filename. We wanted it at the end. And changed to MOS.

There was a lot of renaming that had to be done, but even if I got the names across the versions to be consistent, there would be the “same” family but in different versions. And again, CC does the upgrades for you, so why would I waste my time and the server’s energy uploading the exact same family 4 times?

The conclusion: before uploading to CC, our library needed two things… renaming and deduplication. And did I mention that there was a little over 10,000 files FOR EACH REVIT VERSION YEAR?

What’s In a Name?

Step 0 – declare a freeze on the active library.

Step 1 – make a copy. Obviously, none of this work should be happening on our live library.

Step 2 – I ran a PowerShell script to index all the folders and files to a CSV. No need to review “live” through Windows. It would be easier to compare and analyze the CSV list vs. getting a file list every time.

I mean, it would be easier for Claude.

This is where AI has really gotten me the ROI. The ability to consume gobs of data, look for trends, build plans, sounds very un-sexy but I cannot imagine doing to successfully on my own. I have PowerRename with RegEx, and could make some random scripts, but even with that something would be missed and the more complicated renaming might not be possible. And I would be going blind from staring at the screen.

I explained to Claude the plan and lay of the land:

  • 4 folders with a crap ton of files, probably named almost identically, we wanted them to be named identically
  • Remove any reference to version year
  • Clean up the company abbreviation by changing it to MOS and putting it at the end
  • Some other alignments while we are at it: no underscores, use dashes instead of commas for hierarchy naming, title case across the board, etc.

To be clear, this wasn’t one single run. This was a series of back and forth, but the best part was the conversation of it. And Claude noticing trends that I never would have. And the manual renaming of tens of thousands of files in a matter of minutes.

And Claude did a much better job that I ever would in tracking the renames. I have a series of very large CSVs with “Old Name” and “New Name” columns if I ever need to dig in and see what family is what. (And I also have a record and set of standards to use when I get Claude to rename all the loadable components in my templates!)

Is it the most perfectly consistent named content library? Nope. Is it far more consistent across the version years than before? 100%.

Cleaning Out the Clones

I wanted to minimize the uploads and let CC do the upgrading. That meant removing any file in 25, 26, and 27 that was in 24… THAT WAS NOT UPDATED. We certainly made changes over the years, and some of the files were not what they were in a prior year.

My guess was those changed or added files amounted to less than 10% of the total content.

But, how do we check? How do we know if the file was simply upgraded and not changed or added?

New files were easy. Did that file exist in a prior library year? Nope? Then it’s new and we keep it.

Changed files were actually pretty easy as well. Most of the time, when folks prep the new library year, they copy they old one and run some kind of batch upgrade with Dynamo or a third-party tool. What that means, is that you could easily see a trend in the modified date. Upgrades were done in batches. And if an entire folder or folders of files had modified dates on the same day, usually within 30 minutes of each other, they were part of the batch. They were NOT changed after the upgrade.

I did a manual spot check, and that just confirmed that yes, those files in the same date range were simply upgraded from the previous year.

Claude did a couple of reviews on the modified dates, and we came up with:

  • 2024 (baseline) – 11,745 files
  • 2025 – 577 modified (4.7%) 371 new (3.1%)
  • 2026 – 892 modified (7.4%) 68 new (0.6%)
  • 2027 – 365 modified (3.0%) 13 new (0.1%)

Instead of 50k+ files to upload to CC, our review yielded 14,031 RFAs. A significant difference in uploading, processing, and overhead.

What’s Next

My copy and renamed and cleaned up library is set. Next steps will be to double check that CC is cleaned up, make some collections, and start uploading.

Wish me luck.


If this was helpful, stick around. New entries go out whenever I have something worth saying… which is more often than you’d think. Subscribe below.