prepped for publish

This commit is contained in:
2023-05-23 23:30:15 +10:00
parent 3fe28c6968
commit 6cbe3095af
17 changed files with 500 additions and 100 deletions
+11 -8
View File
@@ -1,5 +1,5 @@
Title: Implmenting Appflow in a Production Datalake
Date: 2023-05-16 20:00
Date: 2023-05-23 20:00
Modified: 2023-05-17 20:00
Category: Data Engineering
Tags: data engineering, Amazon, Managed Services
@@ -27,14 +27,17 @@ This proliferation of API extractors obviously coinccides with the proliferation
This complexity for access is normally coupled with poor documentation, where its a crapshoot as to whether there is an swaggerui, let alone useful API documentation (this is getting better though)
### So why Managed for Extraction?
As you see above when you're extracting data it is so often a crapshoot that writing something bespoke is so incredibally risky that the idea of it gives me hives. I could write a containerised python function for each of my API extractions, or a small batch loader for RDBMS myself and have a small cluster of these things extracting from tables and API endpoints but the though of managing all of that, especially in a 1 man DataOps team is far to overwhelming.
As you see above when you're extracting data it is so often a crapshoot and writing something bespoke is so incrediblly risky that the idea of it gives me hives. I could write a containerised python function for each of my API extractions, or a small batch loader for RDBMS myself and have a small cluster of these things extracting from tables and API endpoints but the thought of managing all of that, especially in a 1 man DataOps team is far to overwhelming.
And Right there is my criteria for choosing a managed server.
And Right there is my criteria for choosing a managed server.
1. Do I want to manage this myself?
2. Is ther any benefit to me manaing this?
2. Is there any benefit to me managing this?
3. Is it more cost effective to have someone else manage it?
Invariably the extraction layer, at least when answering the questions above, gives me the irks and I just decide to run with a simple managed service where I can point at the source and target click go and watch it go brrrrrrrrrrrrr
Invariably, the extraction layer, at least when answering the questions above, gives me the irks and I just decide to run with a simple managed service where I can point at the source and target click go and watch it go brrrrrrrrrrrrr
When you couple ease of use with the relative reliability the value proposition of designing bespoke applications for the extraction task rapidly decreases, at least for me
@@ -42,11 +45,11 @@ And this is why Extraction, at least in systems I design, is more often than not
### AppFlow, The Good, The Bad, The Ugly
Using AppFlow turned out to be a largely simple affair, even in Terraform, Once you have the correct Authentication tokens its more or less select the service you want and then create a "flow" for each endpoint. The complex part is the "Map_All" function for the endpoint. When triggered it automtically create a 1 - 1 mapping for all fields in the endpoint into the target file (in my case parquet) BUT this actually fundamentaly changes the flow you have created and thus causes terraform to shit the bed. THis can be dealt with via a lifecycle rule, but means schema changes in the endpoint could cause issues in the future.
Using AppFlow turned out to be a largely simple affair, even in Terraform, Once you have the correct Authentication tokens its more or less select the service you want and then create a "flow" for each endpoint. The complex part is the "Map_All" function for the endpoint. When triggered it automtically create a 1 - 1 mapping for all fields in the endpoint into the target file (in my case parquet) BUT this actually fundamentaly changes the flow you have created and thus causes terraform to shit the bed. This can be dealt with via a lifecycle rule, but means schema changes in the endpoint could cause issues in the future.
All in All having a Managed Service to managed API endpoint extraction has been great and enabled the expansion of a datalake with no bespoke application code to manage hte extraction of information from API endpoints which has proved to be a massive time and money saver overall
All in All having a Managed Service to manage API endpoint extraction has been great and enabled the expansion of a datalake with no bespoke application code to manage the extraction of information from API endpoints which has proved to be a massive time and money saver overall
I am yet to play with establishing a custom endpoint and it will be interesting to see just hwo much work this is compared with writing the code for a bespoke application... sounds like a good blog post if I get to do it one day.
I am yet to play with establishing a custom endpoint and it will be interesting to see just how much work this is compared with writing the code for a bespoke application... sounds like a good blog post if I get to do it one day.