Every time you download a file from S3, you pay for the traffic you use — the larger the file, the larger the bill.
On average, AWS charges $0.09 per GB, which means a lot of you wouldn't even notice the cost for small files, but things might get expensive for larger ones.
If you have a bunch of large files, for example 5 GB each, you would pay around $0.45 per file to download them.
$0.45 per download is quite expensive; for a lot of businesses it can add up to a huge S3 bill over time.
The interesting thing is that S3 doesn't compress anything for cheaper/faster transfer over the wire, which is strange because it uses HTTP, and HTTP has built-in support for compression.
Anyway, there is another HTTP feature, which we can use as a hack to avoid paying for downloads — for free, without any consequences.
HTTP has an ETag feature, which you may have seen as a 304 Not Modified. The idea, for S3, is that when the client first downloads some file, it stores both the downloaded file itself and an identifier for its version, which S3 sends along with it. On subsequent requests for that same file, it sends that identifier back to the server. If the file hasn't changed, S3 will respond with a 304 Not Modified status and won't charge for the download.
It basically sounds like:
hey, the version I saved locally is X — if it hasn't changed, don't send it again pls, I'll reuse what I already have.
This way, the request that previously cost you ~$0.45 and some time will now cost you nothing and will be much faster, as the file won't be transferred again if it hasn't changed.
You don't need to worry about out-of-date files, as the server checks and responds with a 304 Not Modified only if the version you have in your cache exactly matches the current version of the file.
It seems none of the official SDKs have built-in support for this. You can try it out with capo, a community-driven SDK I am building for Python, with a lot of cool features.