Learn CodeNotes from sessions with Claude Code, my teacher
Lesson 01

Taming a phone's storage from the command line

2026-10-05 · about 3 minutes
command linefileshashingorganisation

What we did#

Over a few sessions we cleaned up a phone with over 100 GB used: deleted large old downloads (on request), moved 171 videos to the SD card, sorted the Download folder by type, and removed exact duplicates. Free space went from 106 GB to over 139 GB.

The tools of the trade#

The phone runs a Linux environment, so the classic command-line tools work.

Command What it does Example
find Search for files by name, type, size, date find Download -iname "*.mkv"
du -sh How much space a folder uses du -sh DCIM/*
df -h How full the whole drive is df -h /internal-storage
mv -n Move without overwriting mv -n file.pdf Documents/PDFs/
md5sum A file's fingerprint (hash) md5sum a.mp4 b.mp4

Why mv -n?

-n means no-clobber: if a file with the same name already exists at the destination, don't overwrite it. When organising hundreds of files, one silent overwrite can destroy something. We logged every clash instead.

Finding true duplicates: hashes#

Two files can have different names but identical content (video.mp4 and video (1).mp4). Comparing names doesn't work; comparing content byte by byte is slow. A hash solves it.

What is a hash?

A hash function (like MD5) turns any file into a short fixed-length fingerprint. Identical files always give identical fingerprints; changing even one byte changes it completely.

The efficient method:

  1. Group files by size first: files of different sizes can't be identical.
  2. Only hash files that share a size with another file.
  3. Same hash = same content.

That turned scanning about 1,750 files into hashing just 306 of them, finding about 2.2 GB of exact duplicates. One surprise: a course's "Section 4" folder was a byte-for-byte copy of "Section 3", so the real Section 4 had never been downloaded.

Check before trusting names

CraigTrade1.mp4 and CraigTrade2.mp4 sounded different but were the same video. maze and maze (1) sounded identical but each had one unique PDF. Content decides, not names.

Organising by type#

We sorted 411 loose files into folders by extension (.pdf → Documents/PDFs, .xlsx → Spreadsheets…). Two subtleties:

  • Files with no extension were identified by their content (file command): many were saved web pages; one was really an Excel file.
  • Hidden files (names starting with a dot, like .env) were deliberately left alone: they usually belong to an app.

Why some files "disappear"#

  • Hidden files: Android hides any name starting with . unless Show hidden system files is on.
  • Search and categories use Android's media index, which doesn't immediately see files created from the command line. Browsing the folder directly always works.

The house rule that came out of this#

After a lot of deleting (always on request), we made it permanent:

  1. Never delete an existing file or folder unless explicitly asked for that specific item.
  2. Always keep the original as a backup before changing anything.
  3. Moves never overwrite.

A delete blocked by the safety check

When a large batch deletion was requested in one go, Claude Code's own safety classifier paused it and asked for clearer confirmation. Listing exactly what would go and getting a "yes, those 27" is the right pattern for anything irreversible.

Key takeaways#

  • du, df and find tell you where the space went.
  • Duplicates are found by size, then hash, never by name.
  • mv -n and logs keep bulk moves safe.
  • Irreversible actions deserve an explicit list and an explicit yes.

Quick quiz#

1. Why group by size before hashing?

Files of different sizes can't be identical, so only same-size files need the (slower) hashing step.

2. What does `-n` do in `mv -n`?

Prevents overwriting a file that already exists at the destination.

Try it yourself#

Find your biggest folders

In a terminal on any computer: du -sh ~/Downloads/* | sort -rh | head -10. That lists the ten biggest items, largest first.