Compare commits
11
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e7794379fb | ||
|
|
40f425eb68 | ||
|
|
23a7ab9674 | ||
|
|
5567aa625d | ||
|
|
8465a992cd | ||
|
|
e636ba5e25 | ||
|
|
2f96fc0b90 | ||
|
|
1074518738 | ||
|
|
d45891ffd9 | ||
|
|
236c1bec1e | ||
|
|
cdb8e28728 |
@@ -1 +0,0 @@
|
|||||||
.venv*
|
|
||||||
@@ -1 +1,2 @@
|
|||||||
pelican[markdown]
|
pelican[markdown]
|
||||||
|
markdown-markup-emoji
|
||||||
|
|||||||
Binary file not shown.
@@ -1,69 +0,0 @@
|
|||||||
Title: CI/CD in Data Engineering
|
|
||||||
Date: 2023-06-15 20:00
|
|
||||||
Modified: 2023-06-15 20:00
|
|
||||||
Category: Data Engineering
|
|
||||||
Tags: data engineering, DBT, Terraform, IAC
|
|
||||||
Slug: CI/CD in Data and Data Infrastructure
|
|
||||||
Authors: Andrew Ridgway
|
|
||||||
Summary: When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture
|
|
||||||
|
|
||||||
Data Engineering has traditionally been considered the bastard step child of work that would have once been considered Administrative in the Tech world. Predominately we write SQL and then deploy that SQL onto one or more Databases. In fact a lot of the traditional methodologies around data almost assume this is the core of how an organistation is managing the majority of it's data. In the last couple of years though there has been a very steady move towards having the Data Engineering workload of SQL move towards Software Engineering techniques. With the popularity of tools like [DBT](https://www.dbtlabs.com) and the latest newcommer on the block, [SQL-MESH](https://www.sqlmesh.com) The oppportunity has started to arise where we can align our Data Engineering workloads with different environments and move much more efficiently towards a Continous Integration and Deployment methodology in our workflows.
|
|
||||||
|
|
||||||
For the Data Engineering space the move to the cloud has been a breath of fresh air (Not so in some other IT disciplines). I am relatively young, so I don't 100% remember but my experience has taught me that there were 3 options here not so long ago
|
|
||||||
|
|
||||||
_Expensive:_
|
|
||||||
|
|
||||||
+ SAS
|
|
||||||
+ SSIS/SSRS
|
|
||||||
+ COGNOS/TM1
|
|
||||||
|
|
||||||
_Rickety:_
|
|
||||||
|
|
||||||
+ Just write stored procedures!
|
|
||||||
+ Startup script on my laptop XD
|
|
||||||
+ "Don't touch that machine over there, No one knows what it does but if it's turned off our financial reports don't work" (This is a third hand story I heard, seriously!)
|
|
||||||
+ "I need to an upgrade to my laptop, Excel needs more than 8GB of RAM"
|
|
||||||
|
|
||||||
_Hard:_
|
|
||||||
|
|
||||||
+ Hadoop
|
|
||||||
+ Spark (hadoop but whatever)
|
|
||||||
+ Python
|
|
||||||
+ R
|
|
||||||
|
|
||||||
_(The reason I've listed them as hard is because self hosting Hadoop/Spark and managing a truckload of python or R scripts, whilst it could have been "cheap" required a team of devs who really really knew what they were doing... so not really cheap and also **really hard**)_
|
|
||||||
|
|
||||||
Then there was getting git behind all the sql scripts and modelling, let alone CI/CD **IF** it existed, it was custom, and bespoke and likely had a single point of failure in the person who knew how `git merge` worked. At least... thats I was told I'm not *that* old ;p.
|
|
||||||
|
|
||||||
These days we are pretty blessed, with democrotisation of clusters and data Infrastructure in the cloud we no longer need a team of sysadmins who know how to tune a cluster to the Nth degree to get the best our of our data workloads (well... we do, but we pay the cloud guys for that!). However, we still need to know about the idiosyncracities of this infrastructure, when it is appropiate to use and how we want to control and maintain the workloads.
|
|
||||||
|
|
||||||
In general when I am designing a system I normally like to break it into 3.
|
|
||||||
|
|
||||||
+ Storage
|
|
||||||
+ Compute
|
|
||||||
+ Code
|
|
||||||
|
|
||||||
*In General* Storage and Compute will be infrastructer related, "Code" is sort of a catch all for my modelling, normally sql, python/r or spark scripts that are used to provide system or business logic, anything really thats going to get data to the end user/analyst/data scientists/annoying person
|
|
||||||
|
|
||||||
Traditionally the compute layer only really had 2 considerations
|
|
||||||
|
|
||||||
+ SQL or Logic engine (normally a flavour of spark(glue) and then something like reshift/athena/trino/bigquery)
|
|
||||||
+ Orchestration Layer (Airflow, Dagster)
|
|
||||||
|
|
||||||
But with the advent of sql engine agnostic Modelling we potentially now need to also consider
|
|
||||||
|
|
||||||
+ Model Compilation
|
|
||||||
|
|
||||||
Now on the surface it seems counterintuitive to seperate the models from the logic layer but lets consider the following scenario
|
|
||||||
|
|
||||||
> Redshift is costing to much and is getting slow, we want to try bigquery
|
|
||||||
> How much investment will it be to change over
|
|
||||||
|
|
||||||
Now, If the entirety of your modelling is stored and deployed to big query direct this would not only involve the investment of spinning up the big query account and either connecting or migrating your data over. You would also need to consider how in the bloody hell you convert all your existing models and workflows over.
|
|
||||||
|
|
||||||
With something like DBT or sqlmesh you change your compilation target and it's done for you. It also means the Data Engineer now doesn't need to necessarily understand the esoteric nature of the target, at least for simple models (which, lets be real, most are).
|
|
||||||
|
|
||||||
BUT, now we have *a lot* of software and infrastructure a simplified common datastack will look something like the below (Assuming ELT, ETL is a bit different but more or less needs the same components)
|
|
||||||
|
|
||||||
<img src="{static}/images/DataStackSimplified.png" width="600" height="295" />
|
|
||||||
|
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
Title: A Cover Letter
|
||||||
|
Date: 2024-02-23 20:00
|
||||||
|
Modified: 2024-03-13 20:00
|
||||||
|
Category: Resume
|
||||||
|
Tags: Cover Letter, Resume
|
||||||
|
Slug: cover-letter
|
||||||
|
Authors: Andrew Ridgway
|
||||||
|
Summary: A Summary of what I've done and Where I'd like to go for prospective Employers
|
||||||
|
|
||||||
|
To whom it may concern
|
||||||
|
|
||||||
|
My name is Andrew Ridgway and I am a Data and Technology professional looking to embark on the next step in my career.
|
||||||
|
|
||||||
|
I have over 10 years’ experience in System and Data Architecture, Data Modelling and Orchestration, Business and Technical Analysis and System and Development Process Design. Most of this has been in developing Cloud architectures and workloads on AWS and GCP Including ML workloads using Sagemaker.
|
||||||
|
|
||||||
|
In my current role I have Proposed, Designed and built the data platform currently used by business. This includes internal and external data products as well as the infrastructure and modelling to support these. This role has seen me liaise with stakeholders of all levels of the business from Analysts in the Customer Experience team right up to C suite executives and preparing material for board members. I understand the complexity of communicating complex system design to different level stakeholders and the complexities of involved in communicating to both technical and less technical employees particularly in relation to data and ML technologies.
|
||||||
|
|
||||||
|
I have also worked as a technical consultant to many businesses and have assisted with the design and implementation of systems for a wide range of industries including financial services, mining and retail. I understand the complexities created by regulation in these environments and understand that this can sometimes necessitate the use of technologies and designs, including legacy systems and designs, I wouldn’t normally use. I also have a passion of designing systems that enable these organisations to realise the benefits of CI/CD on workloads they would not traditionally use this capability. In particular I took a very traditional legacy Data Warehousing team and implemented a solution that meant version control was no longer controlled by a daily copy and paste of folders with dates on major updates. My solution involved establishing guidelines of use of git version control so that this could happen automatically as people committed new code to the core code base. As I have moved into cloud architecture I have made sure to use best practice and ensure everything I build isn’t considered production ready until it is in IAC and deployed through a CI/CD pipeline.
|
||||||
|
|
||||||
|
In a personal capacity I am an avid tech and ML enthusiast. I have designed my own cluster including monitoring and deployment that runs several services that my family uses including chat and DNS and am in the process of designing a “set and forget” system that will allows me to have multi user tenancies on hardware I operate that should enable us to have the niceties of cloud services like email, storage and scheduling with the safety of knowing where that data is stored and exactly how it is used. I also like to design small IoT devices out of Arduino boards allowing me to monitor and control different facets of our house like temperature and light.
|
||||||
|
|
||||||
|
Currently I am working on a project to merge my skill in SQL Modelling and Orchestration with GPT API’s to try and lessen that burden. You can see some of this work in its very early stages here:
|
||||||
|
(gpt-sql-generator)[https://github.com/armistace/gpt-sql-generator]
|
||||||
|
(dbt_sources_generator)[https://github.com/armistace/datahub_dbt_sources_generator]
|
||||||
|
|
||||||
|
|
||||||
|
I look forward to hearing from you soon.
|
||||||
|
|
||||||
|
Sincerely,
|
||||||
|
_________________
|
||||||
|
Andrew Ridgway
|
||||||
|
|
||||||
|
|
||||||
Binary file not shown.
|
Before Width: | Height: | Size: 323 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 146 KiB |
@@ -0,0 +1,105 @@
|
|||||||
|
Title: Metabase and DuckDB
|
||||||
|
Date: 2023-11-15 20:00
|
||||||
|
Modified: 2023-11-15 20:00
|
||||||
|
Category: Business Intelligence
|
||||||
|
Tags: data engineering, Metabase, DuckDB, embedded
|
||||||
|
Slug: metabase-duckdb
|
||||||
|
Authors: Andrew Ridgway
|
||||||
|
Summary: Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible
|
||||||
|
|
||||||
|
Ahhhh [DuckDB](https://duckdb.org/) if you're even partly floating around in the data space you've probably been hearing ALOT about it and it's _"Datawarehouse on your laptop"_ mantra. However, the OTHER application that sometimes gets missed is _"SQLite for OLAP workloads"_ and it was this concept that once I grasped it gave me a very interesting idea.... What if we could take the very pretty Aggregate Layer of our Data(warehouse/LakeHouse/Lake) and put that data right next to presentation layer of the lake, reducing network latency and... hopefully... have presentation reports running over very large workloads in the blink of an eye. It might even be fast enough that it could be deployed and embedded
|
||||||
|
|
||||||
|
However, for this to work we need some form of conatinerised reporting application.... lucky for us there is [Metabase](https://www.metabase.com/) which is a fantastic little reporting application that has an open core. So this got me thinking... Can I put these two applications together and create a Reporting Layer with report embedding capabilities that is deployable in the cluster and has a admin UI accesible over a web page all whilst keeping the data locked to our network?
|
||||||
|
|
||||||
|
### The Beginnings of an Idea
|
||||||
|
Ok so... Big first question. Can Duckdb and Metabase talk? Well... not quite. But first lets take a quick look at the architecture we'll be employing here
|
||||||
|
|
||||||
|
<img alt="Duckdb Architecture" height="auto" width="100%" src="{attach}/images/metabase_duckdb.png">
|
||||||
|
|
||||||
|
But you'll notice this pretty glossed over line, "Connector", that right there is the clincher. So what is this "Connector"?.
|
||||||
|
|
||||||
|
To Deep dive into this would take a whole blog so to give you something to quickly wrap your head around its the glue that will make metabase be able to query your data source. The reality is its a jdbc driver compiled against metabase.
|
||||||
|
|
||||||
|
Thankfully Metabase point you to a [community driver](https://github.com/AlexR2D2/metabase_duckdb_driver) for linking to duckdb ( hopefully it will be brought into metabase proper sooner rather than later )
|
||||||
|
|
||||||
|
Now the release of this driver is still compiled against 0.8 of duckdb and 0.9 is the latest stable but hopefully the [PR](https://github.com/AlexR2D2/metabase_duckdb_driver/pull/19) for this will land very soon giving a good quick way to link to the latest and greatest in duckdb from metabase
|
||||||
|
|
||||||
|
### But How do we get Data?
|
||||||
|
Brilliant, using the recomended DockerFile we can load up a metabase container with the duckdb driver pre built
|
||||||
|
```
|
||||||
|
FROM openjdk:19-buster
|
||||||
|
|
||||||
|
ENV MB_PLUGINS_DIR=/home/plugins/
|
||||||
|
|
||||||
|
ADD https://downloads.metabase.com/v0.46.2/metabase.jar /home
|
||||||
|
ADD https://github.com/AlexR2D2/metabase_duckdb_driver/releases/download/0.1.6/duckdb.metabase-driver.jar /home/plugins/
|
||||||
|
|
||||||
|
RUN chmod 744 /home/plugins/duckdb.metabase-driver.jar
|
||||||
|
|
||||||
|
CMD ["java", "-jar", "/home/metabase.jar"]
|
||||||
|
```
|
||||||
|
|
||||||
|
Great Now the big question. How do we get the data into the damn thing. Interestingly initially when I was designing this I had the thought of leveraging the in memory capabilities of duckdb and pulling in from the parquet on s3 directly as needed, after all the cluster is on AWS so the s3 API requests should be unbelievably fast anyway so why bother with a persistent database?
|
||||||
|
|
||||||
|
Now that we have the default credentials chain it is trivial to call parquet from s3
|
||||||
|
|
||||||
|
```sql
|
||||||
|
SELECT * FROM read_parquet('s3://<bucket>/<file>');
|
||||||
|
```
|
||||||
|
|
||||||
|
However, if you're reading direct off parquet all of a sudden you need to consider the partioning and I also found out that, if the parquet is being actively written to at the time of quering, duckdb has a hissyfit about metadata not matching the query. Needless to say duckdb and streaming parquet are not happy bed fellows (*and frankly were not desined to be so this is ok*). And the idea of trying to explain all this to the run of the mill reporting analyst whom it is my hope is a business sort of person not tech honestly gave me hives.. so I had to make it easier
|
||||||
|
|
||||||
|
The compromise occured to me... the curated layer is only built daily for reporting, and using that, I could create a duckdb file on disk that could be loaded into the metabase container itself.
|
||||||
|
|
||||||
|
With some very simple python as an operation in our orchestrator I had a job that would read direct from our curated parquet and create a duckdb file with it.. without giving away to much the job primarily consisted of this
|
||||||
|
|
||||||
|
```python
|
||||||
|
def duckdb_builder(table):
|
||||||
|
conn = duckdb.connect("curated_duckdb.duckdb")
|
||||||
|
conn.sql(f"CALL load_aws_credentials('{aws_profile}')")
|
||||||
|
#This removes a lot of weirdass ANSI in logs you DO NOT WANT
|
||||||
|
conn.execute("PRAGMA enable_progress_bar=false")
|
||||||
|
log.info(f"Create {table} in duckdb")
|
||||||
|
sql = f"CREATE OR REPLACE TABLE {table} AS SELECT * FROM read_parquet('s3://{curated_bucket}/{table}/*')"
|
||||||
|
conn.sql(sql)
|
||||||
|
log.info(f"{table} Created")
|
||||||
|
```
|
||||||
|
|
||||||
|
And then an upload to an s3 bucket
|
||||||
|
|
||||||
|
This of course necessated a cron job baked in to the metabase container itself to actually pull the duckdb in every morning. After some carefuly analysis of time (because I'm do lazy to implement message queues) I set up a s3 cp job that could be cronned direct from the container itself. This gives us a self updating metabase container pulling with a duckdb backend for client facing reporting right in the interface. AND because of the fact the duckdb is baked right into the container... there are NO associated s3 or dpu costs (merely the cost of running a relatively large container)
|
||||||
|
|
||||||
|
The final Dockerfile looks like this
|
||||||
|
```
|
||||||
|
FROM openjdk:19-buster
|
||||||
|
|
||||||
|
ENV MB_PLUGINS_DIR=/home/plugins/
|
||||||
|
|
||||||
|
ADD https://downloads.metabase.com/v0.47.6/metabase.jar /home
|
||||||
|
ADD duckdb.metabase-driver.jar /home/plugins/
|
||||||
|
|
||||||
|
RUN chmod 744 /home/plugins/duckdb.metabase-driver.jar
|
||||||
|
|
||||||
|
RUN mkdir -p /duckdb_data
|
||||||
|
|
||||||
|
COPY entrypoint.sh /home
|
||||||
|
|
||||||
|
COPY helper_scripts/download_duckdb.py /home
|
||||||
|
|
||||||
|
RUN apt-get update -y && apt-get upgrade -y
|
||||||
|
|
||||||
|
RUN apt-get install python3 python3-pip cron -y
|
||||||
|
|
||||||
|
RUN pip3 install boto3
|
||||||
|
|
||||||
|
RUN crontab -l | { cat; echo "0 */6 * * * python3 /home/helper_scripts/download_duckdb.py"; } | crontab -
|
||||||
|
|
||||||
|
CMD ["bash", "/home/entrypoint.sh"]
|
||||||
|
```
|
||||||
|
|
||||||
|
And there we have it... an in memory containerised reporting solution with blazing fast capability to aggregate and build reports based on curated data direct from the business.. fully automated and deployable via CI/CD, that provides data updates daily.
|
||||||
|
|
||||||
|
Now the embedded part.. which isn't built yet but I'll make sure to update you once we have/if we do because the architecture is very exciting for an embbdedded reporting workflow that is deployable via CI/CD processes to applications. As a little taster I'll point you to the [metabase documentation](https://www.metabase.com/learn/administration/git-based-workflow), the unfortunate thing about it is Metabase *have* hidden this behind the enterprise license.. but I can absolutely see why. If we get to implementing this I'll be sure to update you here on the learnings.
|
||||||
|
|
||||||
|
Until then....
|
||||||
|
|
||||||
@@ -0,0 +1,131 @@
|
|||||||
|
Title: A Resume
|
||||||
|
Date: 2024-02-23 20:00
|
||||||
|
Modified: 2024-03-13 20:00
|
||||||
|
Category: Resume
|
||||||
|
Tags: Cover Letter, Resume
|
||||||
|
Slug: resume
|
||||||
|
Authors: Andrew Ridgway
|
||||||
|
Summary: A Summary of My work Experience
|
||||||
|
|
||||||
|
# OVERVIEW
|
||||||
|
I am a Senior Data Engineer looking to transition my skills to Data and Solution
|
||||||
|
Architecting as well as project management. I have spent the better part of the
|
||||||
|
last decade refining my abilities in taking business requirements and turning
|
||||||
|
those into actionable data engineering, analytics, and software projects with
|
||||||
|
trackable metrics. I believe in agnosticism when it comes to coding languages
|
||||||
|
and have experimented in my own time with many different languages. In my
|
||||||
|
career I have used Python, .NET, PowerShell, TSQL, VB and SAS (multiple
|
||||||
|
products) in an Enterprise capacity. I also have experience using Google Cloud
|
||||||
|
Platform and AWS tools for ETL and data platform development as well as git
|
||||||
|
for version control and deployment using various IAC tools. I have also
|
||||||
|
conducted data analysis and modelling on business metrics to find relationships
|
||||||
|
between both staff and customer behavior and produced actionable
|
||||||
|
recommendations based on the conclusions. In a private context I have also
|
||||||
|
experimented with C, C# and Kotlin I am looking to further my career by taking
|
||||||
|
my passion for data engineering and analysis as well as web and software
|
||||||
|
development and applying it in a strategic context.
|
||||||
|
|
||||||
|
# SKILLS & ABILITIES
|
||||||
|
- Python (scripting, compiling, notebooks – Sagemaker, Jupyter)
|
||||||
|
- git
|
||||||
|
- SAS (Base, EG, VA)
|
||||||
|
- Various Google Cloud Tools (Data Fusion, Compute Engine, Cloud Functions)
|
||||||
|
- Various Amazon Tools (EC2, RDS, Kinesis, Glue, Redshift, Lambda, ECS, ECR, EKS)
|
||||||
|
- Streaming Technologies (Kafka, Hive, Spark Streaming)
|
||||||
|
- Various DB platforms both on Prem and Serverless (MariaDB/MySql,
|
||||||
|
- Postgres/Redshift, SQL Server, RDS/Aurora variants)
|
||||||
|
- Various Microsoft Products (PowerBI, TSQL, Excel, VBA)
|
||||||
|
- Linux Server Administration (cron, bash, systemD)
|
||||||
|
- ETL/ELT Development
|
||||||
|
- Basic Data Modelling (Kimball, SCD Type 2)
|
||||||
|
- IAC (Cloud Formation, Terraform)
|
||||||
|
- Datahub Deployment
|
||||||
|
- Dagster Orchestration Deployments
|
||||||
|
- DBT Modelling and Design Deployments
|
||||||
|
- Containerised and Cloud Driven Data Architecture
|
||||||
|
|
||||||
|
# EXPERIENCE
|
||||||
|
## Cloud Data Architect
|
||||||
|
### _Redeye Apps_
|
||||||
|
#### _May 2022 - Present_
|
||||||
|
- Greenfields Research, Design and Deployment of S3 datalake (Parquet)
|
||||||
|
- AWS DMS, S3, Athena, Glue
|
||||||
|
- Research Design and Deployment of Catalog (Datahub)
|
||||||
|
- Design of Data Governance Process (Datahub driven)
|
||||||
|
- Research Design and Deployment of Orchestration and Modelling for Transforms (Dagster/DBT into Mesos)
|
||||||
|
- CI/CD design and deployment of modelling and orchestration using Gitlab
|
||||||
|
- Research, Design and Deployment of ML Ops Dev pipelines anddeployment strategy
|
||||||
|
- Design of ETL/Pipelines (DBT)
|
||||||
|
- Design of Customer Facing Data Products and deployment methodologies (Fully automated via Kakfa/Dagster/DBT)
|
||||||
|
|
||||||
|
## Data Engineer,
|
||||||
|
### _TechConnect IT Solutions_
|
||||||
|
#### _August 2021 – May 2022_
|
||||||
|
- Design of Cloud Data Batch ETL solutions using Python (Glue)
|
||||||
|
- Design of Cloud Data Streaming ETL solution using Python (Kinesis)
|
||||||
|
- Solve complex client business problems using software to join and transform data from DB’s, Web API’s, Application API’s and System logs
|
||||||
|
- Build CI/CD pipelines to ensure smooth deployments (Bitbucket, gitlab)
|
||||||
|
- Apply Prebuilt ML models to software solutions (Sagemaker)
|
||||||
|
- Assist with the architecting of Containerisation solutions (Docker, ECS, ECR)
|
||||||
|
- API testing and development (gRPC, Rest)
|
||||||
|
|
||||||
|
## Enterprise Data Warehouse Developer
|
||||||
|
### _Auto and General Insurance_
|
||||||
|
#### _August 2019 - August 2021_
|
||||||
|
- ETL development of CRM, WFP, Outbound Dialer, Inbound switch in Google Cloud, SAS, TSQL
|
||||||
|
- Bringing new data to the business to analyse for new insights
|
||||||
|
- Redeveloped Version Control and brought git to the data team
|
||||||
|
- Introduced python for API enablement in the Enterprise Data Warehouse
|
||||||
|
- Partnering with the business to focus data project on actual need and translating into technical requirements
|
||||||
|
|
||||||
|
## Business Analyst
|
||||||
|
### _Auto and General Insurance_
|
||||||
|
#### _January 2018 - August 2019_
|
||||||
|
- Automate Service Performance Reporting using PowerShell/VBA/SAS
|
||||||
|
- Learn and leverage SAS EG and VA to streamline Microsoft Excel Reporting
|
||||||
|
- Identify and develop data pipelines to source data from multiple sources easily and collate into a single source to identify relationships and trends
|
||||||
|
- Technologies used include VBA, PowerShell, SQL, Web API’s, SAS
|
||||||
|
- Where SAS is inappropriate use VBA to automate processes in Microsoft Access and Excel
|
||||||
|
- Gather Requirements to build meaningful reporting solutions
|
||||||
|
- Provide meaningful analysis on business performance and provide relevant presentations and reports to senior stakeholders.
|
||||||
|
|
||||||
|
## Forecasting and Capacity Analyst
|
||||||
|
### _Auto and General Insurance_
|
||||||
|
#### _January 2017 – January 2018_
|
||||||
|
- Develop the outbound forecasting model for the Auto and General sales call center by analysing the relationship between customer decisions and workload drivers
|
||||||
|
- This includes the complete data pipeline for the model from identifying and sourcing data, building the reporting and analysing the data and associated drivers.
|
||||||
|
- Forecast inbound workload requirements for the Auto and General sales call center using time series analysis
|
||||||
|
- Learn and leverage the Aspect Workforce Management System to ensure efficiency of forecast generation
|
||||||
|
- Learn and leverage the capabilities of SAS Enterprise Guide to improve accuracy
|
||||||
|
- Liaise with people across the business to ensure meaningful, accurate analysis is provided to senior stakeholders
|
||||||
|
- Analyse monthly, weekly and intraday requirements and ensure forecast is accurately predicting workload for breaks, meetings and Leave
|
||||||
|
|
||||||
|
## Senior HR Performance Analyst
|
||||||
|
### _Queensland Department of Justice and Attorney General_
|
||||||
|
#### _June 2016 - January 2017_
|
||||||
|
- Harmonise various systems to develop a unified workforce reporting and analysis framework with appropriate metrics
|
||||||
|
- Use VBA to automate regular reporting in Microsoft Access and Excel
|
||||||
|
- Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives
|
||||||
|
|
||||||
|
## Workforce Business Analyst
|
||||||
|
### _Queensland Department of Justice and Attorney General_
|
||||||
|
#### _July 2015 – June 2016_
|
||||||
|
- Develop and refine current workforce analysis techniques and databases
|
||||||
|
- Use VBA to automate regular reporting in Microsoft Access and Excel
|
||||||
|
- Act as liaison between shared service providers and executives and facilitate communication during the implementation of a payroll leave audit
|
||||||
|
- Gather reporting requirements from various business areas and produce ad-hoc and regular reports as required
|
||||||
|
- Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives
|
||||||
|
|
||||||
|
# EDUCATION
|
||||||
|
- 2011 Bachelor of Business Management, University of Queensland
|
||||||
|
- 2008 Bachelor of Arts, University of Queensland
|
||||||
|
|
||||||
|
# REFERENCES
|
||||||
|
- Anthony Stiller Lead Developer, Data warehousing, Queensland Health
|
||||||
|
|
||||||
|
_0428 038 031_
|
||||||
|
|
||||||
|
- Jaime Brian Head of Cloud Ninjas, TechConnect
|
||||||
|
|
||||||
|
_0422 012 17_
|
||||||
|
|
||||||
@@ -16,6 +16,5 @@ TWITTER_URL = 'https://twitter.com/ar17787'
|
|||||||
FACEBOOK_URL = 'https://facebook.com/ar17787'
|
FACEBOOK_URL = 'https://facebook.com/ar17787'
|
||||||
|
|
||||||
DEFAULT_PAGINATION = 10
|
DEFAULT_PAGINATION = 10
|
||||||
|
|
||||||
# Uncomment following line if you want document-relative URLs when developing
|
# Uncomment following line if you want document-relative URLs when developing
|
||||||
#RELATIVE_URLS = True
|
#RELATIVE_URLS = True
|
||||||
|
|||||||
@@ -82,8 +82,12 @@
|
|||||||
<div class="row">
|
<div class="row">
|
||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<dl>
|
<dl>
|
||||||
<dt>Thu 15 June 2023</dt>
|
<dt>Fri 23 February 2024</dt>
|
||||||
<dd><a href="http://localhost:8000/CI/CD in Data and Data Infrastructure.html">CI/CD in Data Engineering</a></dd>
|
<dd><a href="http://localhost:8000/cover-letter.html">A Cover Letter</a></dd>
|
||||||
|
<dt>Fri 23 February 2024</dt>
|
||||||
|
<dd><a href="http://localhost:8000/resume.html">A Resume</a></dd>
|
||||||
|
<dt>Wed 15 November 2023</dt>
|
||||||
|
<dd><a href="http://localhost:8000/metabase-duckdb.html">Metabase and DuckDB</a></dd>
|
||||||
<dt>Tue 23 May 2023</dt>
|
<dt>Tue 23 May 2023</dt>
|
||||||
<dd><a href="http://localhost:8000/appflow-production.html">Implmenting Appflow in a Production Datalake</a></dd>
|
<dd><a href="http://localhost:8000/appflow-production.html">Implmenting Appflow in a Production Datalake</a></dd>
|
||||||
<dt>Wed 10 May 2023</dt>
|
<dt>Wed 10 May 2023</dt>
|
||||||
|
|||||||
@@ -82,15 +82,41 @@
|
|||||||
<div class="row">
|
<div class="row">
|
||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<div class="post-preview">
|
<div class="post-preview">
|
||||||
<a href="http://localhost:8000/CI/CD in Data and Data Infrastructure.html" rel="bookmark" title="Permalink to CI/CD in Data Engineering">
|
<a href="http://localhost:8000/cover-letter.html" rel="bookmark" title="Permalink to A Cover Letter">
|
||||||
<h2 class="post-title">
|
<h2 class="post-title">
|
||||||
CI/CD in Data Engineering
|
A Cover Letter
|
||||||
</h2>
|
</h2>
|
||||||
</a>
|
</a>
|
||||||
<p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p>
|
<p>A Summary of what I've done and Where I'd like to go for prospective Employers</p>
|
||||||
<p class="post-meta">Posted by
|
<p class="post-meta">Posted by
|
||||||
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
on Thu 15 June 2023
|
on Fri 23 February 2024
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/resume.html" rel="bookmark" title="Permalink to A Resume">
|
||||||
|
<h2 class="post-title">
|
||||||
|
A Resume
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>A Summary of My work Experience</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Fri 23 February 2024
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/metabase-duckdb.html" rel="bookmark" title="Permalink to Metabase and DuckDB">
|
||||||
|
<h2 class="post-title">
|
||||||
|
Metabase and DuckDB
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Wed 15 November 2023
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
<hr>
|
<hr>
|
||||||
|
|||||||
@@ -84,7 +84,7 @@
|
|||||||
<div class="post-preview">
|
<div class="post-preview">
|
||||||
<a href="http://localhost:8000/author/andrew-ridgway.html" rel="bookmark">
|
<a href="http://localhost:8000/author/andrew-ridgway.html" rel="bookmark">
|
||||||
<h2 class="post-title">
|
<h2 class="post-title">
|
||||||
Andrew Ridgway (3)
|
Andrew Ridgway (5)
|
||||||
</h2>
|
</h2>
|
||||||
</a>
|
</a>
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
@@ -82,7 +82,9 @@
|
|||||||
<div class="row">
|
<div class="row">
|
||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<ul>
|
<ul>
|
||||||
|
<li><a href="http://localhost:8000/category/business-intelligence.html">Business Intelligence</a></li>
|
||||||
<li><a href="http://localhost:8000/category/data-engineering.html">Data Engineering</a></li>
|
<li><a href="http://localhost:8000/category/data-engineering.html">Data Engineering</a></li>
|
||||||
|
<li><a href="http://localhost:8000/category/resume.html">Resume</a></li>
|
||||||
</ul>
|
</ul>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
@@ -0,0 +1,165 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<meta name="description" content="">
|
||||||
|
<meta name="author" content="">
|
||||||
|
|
||||||
|
<title>Andrew Ridgway's Blog - Articles in the Business Intelligence category</title>
|
||||||
|
|
||||||
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
|
<link href="http://localhost:8000/feeds/business-intelligence.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
|
<!-- Bootstrap Core CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/clean-blog.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Code highlight color scheme -->
|
||||||
|
<link href="http://localhost:8000/theme/css/code_blocks/tomorrow.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom Fonts -->
|
||||||
|
<link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'>
|
||||||
|
|
||||||
|
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
|
||||||
|
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
|
||||||
|
<!--[if lt IE 9]>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
|
||||||
|
<![endif]-->
|
||||||
|
|
||||||
|
<meta property="og:locale" content="en">
|
||||||
|
<meta property="og:site_name" content="Andrew Ridgway's Blog">
|
||||||
|
</head>
|
||||||
|
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<!-- Navigation -->
|
||||||
|
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
|
||||||
|
<div class="container-fluid">
|
||||||
|
<!-- Brand and toggle get grouped for better mobile display -->
|
||||||
|
<div class="navbar-header page-scroll">
|
||||||
|
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
|
||||||
|
<span class="sr-only">Toggle navigation</span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
</button>
|
||||||
|
<a class="navbar-brand" href="http://localhost:8000/">Andrew Ridgway's Blog</a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Collect the nav links, forms, and other content for toggling -->
|
||||||
|
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
|
||||||
|
<ul class="nav navbar-nav navbar-right">
|
||||||
|
|
||||||
|
</ul>
|
||||||
|
</div>
|
||||||
|
<!-- /.navbar-collapse -->
|
||||||
|
</div>
|
||||||
|
<!-- /.container -->
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<!-- Page Header -->
|
||||||
|
<header class="intro-header" style="background-image: url('https://wallpaperaccess.com/full/3239444.jpg')">
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-heading">
|
||||||
|
<h1>Articles in the Business Intelligence category</h1>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- Main Content -->
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/metabase-duckdb.html" rel="bookmark" title="Permalink to Metabase and DuckDB">
|
||||||
|
<h2 class="post-title">
|
||||||
|
Metabase and DuckDB
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Wed 15 November 2023
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Pager -->
|
||||||
|
<ul class="pager">
|
||||||
|
<li class="next">
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
Page 1 / 1
|
||||||
|
<hr>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Footer -->
|
||||||
|
<footer>
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<p>
|
||||||
|
<script type="text/javascript" src="https://sessionize.com/api/speaker/sessions/83c5d14a-bd19-46b4-8335-0ac8358ac46d/0x0x91929ax">
|
||||||
|
</script>
|
||||||
|
</p>
|
||||||
|
<ul class="list-inline text-center">
|
||||||
|
<li>
|
||||||
|
<a href="https://twitter.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-twitter fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://facebook.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-facebook fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://github.com/armistace">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-github fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
<p class="copyright text-muted">Blog powered by <a href="http://getpelican.com">Pelican</a>,
|
||||||
|
which takes great advantage of <a href="http://python.org">Python</a>.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</footer>
|
||||||
|
|
||||||
|
<!-- jQuery -->
|
||||||
|
<script src="http://localhost:8000/theme/js/jquery.js"></script>
|
||||||
|
|
||||||
|
<!-- Bootstrap Core JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/bootstrap.min.js"></script>
|
||||||
|
|
||||||
|
<!-- Custom Theme JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/clean-blog.min.js"></script>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
|
||||||
|
</html>
|
||||||
@@ -0,0 +1,165 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<meta name="description" content="">
|
||||||
|
<meta name="author" content="">
|
||||||
|
|
||||||
|
<title>Andrew Ridgway's Blog - Articles in the Data Analytics category</title>
|
||||||
|
|
||||||
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
|
<link href="http://localhost:8000/feeds/data-analytics.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
|
<!-- Bootstrap Core CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/clean-blog.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Code highlight color scheme -->
|
||||||
|
<link href="http://localhost:8000/theme/css/code_blocks/tomorrow.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom Fonts -->
|
||||||
|
<link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'>
|
||||||
|
|
||||||
|
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
|
||||||
|
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
|
||||||
|
<!--[if lt IE 9]>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
|
||||||
|
<![endif]-->
|
||||||
|
|
||||||
|
<meta property="og:locale" content="en">
|
||||||
|
<meta property="og:site_name" content="Andrew Ridgway's Blog">
|
||||||
|
</head>
|
||||||
|
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<!-- Navigation -->
|
||||||
|
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
|
||||||
|
<div class="container-fluid">
|
||||||
|
<!-- Brand and toggle get grouped for better mobile display -->
|
||||||
|
<div class="navbar-header page-scroll">
|
||||||
|
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
|
||||||
|
<span class="sr-only">Toggle navigation</span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
</button>
|
||||||
|
<a class="navbar-brand" href="http://localhost:8000/">Andrew Ridgway's Blog</a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Collect the nav links, forms, and other content for toggling -->
|
||||||
|
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
|
||||||
|
<ul class="nav navbar-nav navbar-right">
|
||||||
|
|
||||||
|
</ul>
|
||||||
|
</div>
|
||||||
|
<!-- /.navbar-collapse -->
|
||||||
|
</div>
|
||||||
|
<!-- /.container -->
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<!-- Page Header -->
|
||||||
|
<header class="intro-header" style="background-image: url('https://wallpaperaccess.com/full/3239444.jpg')">
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-heading">
|
||||||
|
<h1>Articles in the Data Analytics category</h1>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- Main Content -->
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/notebook-or-bi.html" rel="bookmark" title="Permalink to Notebook or BI, What is the most appropiate communication medium">
|
||||||
|
<h2 class="post-title">
|
||||||
|
Notebook or BI, What is the most appropiate communication medium
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>When is a notebook enough or when do we need a dashboard</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Thu 13 July 2023
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Pager -->
|
||||||
|
<ul class="pager">
|
||||||
|
<li class="next">
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
Page 1 / 1
|
||||||
|
<hr>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Footer -->
|
||||||
|
<footer>
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<p>
|
||||||
|
<script type="text/javascript" src="https://sessionize.com/api/speaker/sessions/83c5d14a-bd19-46b4-8335-0ac8358ac46d/0x0x91929ax">
|
||||||
|
</script>
|
||||||
|
</p>
|
||||||
|
<ul class="list-inline text-center">
|
||||||
|
<li>
|
||||||
|
<a href="https://twitter.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-twitter fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://facebook.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-facebook fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://github.com/armistace">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-github fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
<p class="copyright text-muted">Blog powered by <a href="http://getpelican.com">Pelican</a>,
|
||||||
|
which takes great advantage of <a href="http://python.org">Python</a>.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</footer>
|
||||||
|
|
||||||
|
<!-- jQuery -->
|
||||||
|
<script src="http://localhost:8000/theme/js/jquery.js"></script>
|
||||||
|
|
||||||
|
<!-- Bootstrap Core JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/bootstrap.min.js"></script>
|
||||||
|
|
||||||
|
<!-- Custom Theme JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/clean-blog.min.js"></script>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
|
||||||
|
</html>
|
||||||
@@ -82,19 +82,6 @@
|
|||||||
<div class="container">
|
<div class="container">
|
||||||
<div class="row">
|
<div class="row">
|
||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<div class="post-preview">
|
|
||||||
<a href="http://localhost:8000/CI/CD in Data and Data Infrastructure.html" rel="bookmark" title="Permalink to CI/CD in Data Engineering">
|
|
||||||
<h2 class="post-title">
|
|
||||||
CI/CD in Data Engineering
|
|
||||||
</h2>
|
|
||||||
</a>
|
|
||||||
<p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p>
|
|
||||||
<p class="post-meta">Posted by
|
|
||||||
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
|
||||||
on Thu 15 June 2023
|
|
||||||
</p>
|
|
||||||
</div>
|
|
||||||
<hr>
|
|
||||||
<div class="post-preview">
|
<div class="post-preview">
|
||||||
<a href="http://localhost:8000/appflow-production.html" rel="bookmark" title="Permalink to Implmenting Appflow in a Production Datalake">
|
<a href="http://localhost:8000/appflow-production.html" rel="bookmark" title="Permalink to Implmenting Appflow in a Production Datalake">
|
||||||
<h2 class="post-title">
|
<h2 class="post-title">
|
||||||
|
|||||||
@@ -0,0 +1,178 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<meta name="description" content="">
|
||||||
|
<meta name="author" content="">
|
||||||
|
|
||||||
|
<title>Andrew Ridgway's Blog - Articles in the Resume category</title>
|
||||||
|
|
||||||
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
|
<link href="http://localhost:8000/feeds/resume.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
|
<!-- Bootstrap Core CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/clean-blog.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Code highlight color scheme -->
|
||||||
|
<link href="http://localhost:8000/theme/css/code_blocks/tomorrow.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom Fonts -->
|
||||||
|
<link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'>
|
||||||
|
|
||||||
|
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
|
||||||
|
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
|
||||||
|
<!--[if lt IE 9]>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
|
||||||
|
<![endif]-->
|
||||||
|
|
||||||
|
<meta property="og:locale" content="en">
|
||||||
|
<meta property="og:site_name" content="Andrew Ridgway's Blog">
|
||||||
|
</head>
|
||||||
|
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<!-- Navigation -->
|
||||||
|
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
|
||||||
|
<div class="container-fluid">
|
||||||
|
<!-- Brand and toggle get grouped for better mobile display -->
|
||||||
|
<div class="navbar-header page-scroll">
|
||||||
|
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
|
||||||
|
<span class="sr-only">Toggle navigation</span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
</button>
|
||||||
|
<a class="navbar-brand" href="http://localhost:8000/">Andrew Ridgway's Blog</a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Collect the nav links, forms, and other content for toggling -->
|
||||||
|
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
|
||||||
|
<ul class="nav navbar-nav navbar-right">
|
||||||
|
|
||||||
|
</ul>
|
||||||
|
</div>
|
||||||
|
<!-- /.navbar-collapse -->
|
||||||
|
</div>
|
||||||
|
<!-- /.container -->
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<!-- Page Header -->
|
||||||
|
<header class="intro-header" style="background-image: url('https://wallpaperaccess.com/full/3239444.jpg')">
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-heading">
|
||||||
|
<h1>Articles in the Resume category</h1>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- Main Content -->
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/cover-letter.html" rel="bookmark" title="Permalink to A Cover Letter">
|
||||||
|
<h2 class="post-title">
|
||||||
|
A Cover Letter
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>A Summary of what I've done and Where I'd like to go for prospective Employers</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Fri 23 February 2024
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/resume.html" rel="bookmark" title="Permalink to A Resume">
|
||||||
|
<h2 class="post-title">
|
||||||
|
A Resume
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>A Summary of My work Experience</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Fri 23 February 2024
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Pager -->
|
||||||
|
<ul class="pager">
|
||||||
|
<li class="next">
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
Page 1 / 1
|
||||||
|
<hr>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Footer -->
|
||||||
|
<footer>
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<p>
|
||||||
|
<script type="text/javascript" src="https://sessionize.com/api/speaker/sessions/83c5d14a-bd19-46b4-8335-0ac8358ac46d/0x0x91929ax">
|
||||||
|
</script>
|
||||||
|
</p>
|
||||||
|
<ul class="list-inline text-center">
|
||||||
|
<li>
|
||||||
|
<a href="https://twitter.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-twitter fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://facebook.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-facebook fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://github.com/armistace">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-github fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
<p class="copyright text-muted">Blog powered by <a href="http://getpelican.com">Pelican</a>,
|
||||||
|
which takes great advantage of <a href="http://python.org">Python</a>.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</footer>
|
||||||
|
|
||||||
|
<!-- jQuery -->
|
||||||
|
<script src="http://localhost:8000/theme/js/jquery.js"></script>
|
||||||
|
|
||||||
|
<!-- Bootstrap Core JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/bootstrap.min.js"></script>
|
||||||
|
|
||||||
|
<!-- Custom Theme JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/clean-blog.min.js"></script>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
|
||||||
|
</html>
|
||||||
@@ -0,0 +1,183 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<meta name="description" content="">
|
||||||
|
<meta name="author" content="">
|
||||||
|
|
||||||
|
<title>Andrew Ridgway's Blog</title>
|
||||||
|
|
||||||
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
|
<link href="http://localhost:8000/feeds/resume.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
|
<!-- Bootstrap Core CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/clean-blog.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Code highlight color scheme -->
|
||||||
|
<link href="http://localhost:8000/theme/css/code_blocks/tomorrow.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom Fonts -->
|
||||||
|
<link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'>
|
||||||
|
|
||||||
|
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
|
||||||
|
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
|
||||||
|
<!--[if lt IE 9]>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
|
||||||
|
<![endif]-->
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
<meta name="tags" contents="Cover Letter" />
|
||||||
|
<meta name="tags" contents="Resume" />
|
||||||
|
|
||||||
|
|
||||||
|
<meta property="og:locale" content="en">
|
||||||
|
<meta property="og:site_name" content="Andrew Ridgway's Blog">
|
||||||
|
|
||||||
|
<meta property="og:type" content="article">
|
||||||
|
<meta property="article:author" content="">
|
||||||
|
<meta property="og:url" content="http://localhost:8000/cover-letter.html">
|
||||||
|
<meta property="og:title" content="A Cover Letter">
|
||||||
|
<meta property="og:description" content="">
|
||||||
|
<meta property="og:image" content="http://localhost:8000/">
|
||||||
|
<meta property="article:published_time" content="2024-02-23 20:00:00+10:00">
|
||||||
|
</head>
|
||||||
|
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<!-- Navigation -->
|
||||||
|
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
|
||||||
|
<div class="container-fluid">
|
||||||
|
<!-- Brand and toggle get grouped for better mobile display -->
|
||||||
|
<div class="navbar-header page-scroll">
|
||||||
|
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
|
||||||
|
<span class="sr-only">Toggle navigation</span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
</button>
|
||||||
|
<a class="navbar-brand" href="http://localhost:8000/">Andrew Ridgway's Blog</a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Collect the nav links, forms, and other content for toggling -->
|
||||||
|
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
|
||||||
|
<ul class="nav navbar-nav navbar-right">
|
||||||
|
|
||||||
|
</ul>
|
||||||
|
</div>
|
||||||
|
<!-- /.navbar-collapse -->
|
||||||
|
</div>
|
||||||
|
<!-- /.container -->
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<!-- Page Header -->
|
||||||
|
<header class="intro-header" style="background-image: url('http://localhost:8000/theme/images/post-bg.jpg')">
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-heading">
|
||||||
|
<h1>A Cover Letter</h1>
|
||||||
|
<span class="meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Fri 23 February 2024
|
||||||
|
</span>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- Main Content -->
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<!-- Post Content -->
|
||||||
|
<article>
|
||||||
|
<p>To whom it may concern</p>
|
||||||
|
<p>My name is Andrew Ridgway and I am a Data and Technology professional looking to embark on the next step in my career.</p>
|
||||||
|
<p>I have over 10 years’ experience in System and Data Architecture, Data Modelling and Orchestration, Business and Technical Analysis and System and Development Process Design. Most of this has been in developing Cloud architectures and workloads on AWS and GCP Including ML workloads using Sagemaker. </p>
|
||||||
|
<p>In my current role I have Proposed, Designed and built the data platform currently used by business. This includes internal and external data products as well as the infrastructure and modelling to support these. This role has seen me liaise with stakeholders of all levels of the business from Analysts in the Customer Experience team right up to C suite executives and preparing material for board members. I understand the complexity of communicating complex system design to different level stakeholders and the complexities of involved in communicating to both technical and less technical employees particularly in relation to data and ML technologies. </p>
|
||||||
|
<p>I have also worked as a technical consultant to many businesses and have assisted with the design and implementation of systems for a wide range of industries including financial services, mining and retail. I understand the complexities created by regulation in these environments and understand that this can sometimes necessitate the use of technologies and designs, including legacy systems and designs, I wouldn’t normally use. I also have a passion of designing systems that enable these organisations to realise the benefits of CI/CD on workloads they would not traditionally use this capability. In particular I took a very traditional legacy Data Warehousing team and implemented a solution that meant version control was no longer controlled by a daily copy and paste of folders with dates on major updates. My solution involved establishing guidelines of use of git version control so that this could happen automatically as people committed new code to the core code base. As I have moved into cloud architecture I have made sure to use best practice and ensure everything I build isn’t considered production ready until it is in IAC and deployed through a CI/CD pipeline.</p>
|
||||||
|
<p>In a personal capacity I am an avid tech and ML enthusiast. I have designed my own cluster including monitoring and deployment that runs several services that my family uses including chat and DNS and am in the process of designing a “set and forget” system that will allows me to have multi user tenancies on hardware I operate that should enable us to have the niceties of cloud services like email, storage and scheduling with the safety of knowing where that data is stored and exactly how it is used. I also like to design small IoT devices out of Arduino boards allowing me to monitor and control different facets of our house like temperature and light. </p>
|
||||||
|
<p>Currently I am working on a project to merge my skill in SQL Modelling and Orchestration with GPT API’s to try and lessen that burden. You can see some of this work in its very early stages here:
|
||||||
|
(gpt-sql-generator)[https://github.com/armistace/gpt-sql-generator]
|
||||||
|
(dbt_sources_generator)[https://github.com/armistace/datahub_dbt_sources_generator]</p>
|
||||||
|
<p>I look forward to hearing from you soon.</p>
|
||||||
|
<p>Sincerely,</p>
|
||||||
|
<hr>
|
||||||
|
<p>Andrew Ridgway</p>
|
||||||
|
</article>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Footer -->
|
||||||
|
<footer>
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<p>
|
||||||
|
<script type="text/javascript" src="https://sessionize.com/api/speaker/sessions/83c5d14a-bd19-46b4-8335-0ac8358ac46d/0x0x91929ax">
|
||||||
|
</script>
|
||||||
|
</p>
|
||||||
|
<ul class="list-inline text-center">
|
||||||
|
<li>
|
||||||
|
<a href="https://twitter.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-twitter fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://facebook.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-facebook fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://github.com/armistace">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-github fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
<p class="copyright text-muted">Blog powered by <a href="http://getpelican.com">Pelican</a>,
|
||||||
|
which takes great advantage of <a href="http://python.org">Python</a>.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</footer>
|
||||||
|
|
||||||
|
<!-- jQuery -->
|
||||||
|
<script src="http://localhost:8000/theme/js/jquery.js"></script>
|
||||||
|
|
||||||
|
<!-- Bootstrap Core JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/bootstrap.min.js"></script>
|
||||||
|
|
||||||
|
<!-- Custom Theme JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/clean-blog.min.js"></script>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
|
||||||
|
</html>
|
||||||
@@ -1,54 +1,219 @@
|
|||||||
<?xml version="1.0" encoding="utf-8"?>
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/all-en.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-06-15T20:00:00+10:00</updated><entry><title>CI/CD in Data Engineering</title><link href="http://localhost:8000/CI/CD%20in%20Data%20and%20Data%20Infrastructure.html" rel="alternate"></link><published>2023-06-15T20:00:00+10:00</published><updated>2023-06-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-06-15:/CI/CD in Data and Data Infrastructure.html</id><summary type="html"><p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p></summary><content type="html"><p>Data Engineering has traditionally been considered the bastard step child of work that would have once been considered Administrative in the Tech world. Predominately we write SQL and then deploy that SQL onto one or more Databases. In fact a lot of the traditional methodologies around data almost assume this is the core of how an organistation is managing the majority of it's data. In the last couple of years though there has been a very steady move towards having the Data Engineering workload of SQL move towards Software Engineering techniques. With the popularity of tools like <a href="https://www.dbtlabs.com">DBT</a> and the latest newcommer on the block, <a href="https://www.sqlmesh.com">SQL-MESH</a> The oppportunity has started to arise where we can align our Data Engineering workloads with different environments and move much more efficiently towards a Continous Integration and Deployment methodology in our workflows. </p>
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/all-en.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2024-03-13T20:00:00+10:00</updated><entry><title>A Cover Letter</title><link href="http://localhost:8000/cover-letter.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/cover-letter.html</id><summary type="html"><p>A Summary of what I've done and Where I'd like to go for prospective Employers</p></summary><content type="html"><p>To whom it may concern</p>
|
||||||
<p>For the Data Engineering space the move to the cloud has been a breath of fresh air (Not so in some other IT disciplines). I am relatively young, so I don't 100% remember but my experience has taught me that there were 3 options here not so long ago</p>
|
<p>My name is Andrew Ridgway and I am a Data and Technology professional looking to embark on the next step in my career.</p>
|
||||||
<p><em>Expensive:</em></p>
|
<p>I have over 10 years’ experience in System and Data Architecture, Data Modelling and Orchestration, Business and Technical Analysis and System and Development Process Design. Most of this has been in developing Cloud architectures and workloads on AWS and GCP Including ML workloads using Sagemaker. </p>
|
||||||
|
<p>In my current role I have Proposed, Designed and built the data platform currently used by business. This includes internal and external data products as well as the infrastructure and modelling to support these. This role has seen me liaise with stakeholders of all levels of the business from Analysts in the Customer Experience team right up to C suite executives and preparing material for board members. I understand the complexity of communicating complex system design to different level stakeholders and the complexities of involved in communicating to both technical and less technical employees particularly in relation to data and ML technologies. </p>
|
||||||
|
<p>I have also worked as a technical consultant to many businesses and have assisted with the design and implementation of systems for a wide range of industries including financial services, mining and retail. I understand the complexities created by regulation in these environments and understand that this can sometimes necessitate the use of technologies and designs, including legacy systems and designs, I wouldn’t normally use. I also have a passion of designing systems that enable these organisations to realise the benefits of CI/CD on workloads they would not traditionally use this capability. In particular I took a very traditional legacy Data Warehousing team and implemented a solution that meant version control was no longer controlled by a daily copy and paste of folders with dates on major updates. My solution involved establishing guidelines of use of git version control so that this could happen automatically as people committed new code to the core code base. As I have moved into cloud architecture I have made sure to use best practice and ensure everything I build isn’t considered production ready until it is in IAC and deployed through a CI/CD pipeline.</p>
|
||||||
|
<p>In a personal capacity I am an avid tech and ML enthusiast. I have designed my own cluster including monitoring and deployment that runs several services that my family uses including chat and DNS and am in the process of designing a “set and forget” system that will allows me to have multi user tenancies on hardware I operate that should enable us to have the niceties of cloud services like email, storage and scheduling with the safety of knowing where that data is stored and exactly how it is used. I also like to design small IoT devices out of Arduino boards allowing me to monitor and control different facets of our house like temperature and light. </p>
|
||||||
|
<p>Currently I am working on a project to merge my skill in SQL Modelling and Orchestration with GPT API’s to try and lessen that burden. You can see some of this work in its very early stages here:
|
||||||
|
(gpt-sql-generator)[https://github.com/armistace/gpt-sql-generator]
|
||||||
|
(dbt_sources_generator)[https://github.com/armistace/datahub_dbt_sources_generator]</p>
|
||||||
|
<p>I look forward to hearing from you soon.</p>
|
||||||
|
<p>Sincerely,</p>
|
||||||
|
<hr>
|
||||||
|
<p>Andrew Ridgway</p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>A Resume</title><link href="http://localhost:8000/resume.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/resume.html</id><summary type="html"><p>A Summary of My work Experience</p></summary><content type="html"><h1>OVERVIEW</h1>
|
||||||
|
<p>I am a Senior Data Engineer looking to transition my skills to Data and Solution
|
||||||
|
Architecting as well as project management. I have spent the better part of the
|
||||||
|
last decade refining my abilities in taking business requirements and turning
|
||||||
|
those into actionable data engineering, analytics, and software projects with
|
||||||
|
trackable metrics. I believe in agnosticism when it comes to coding languages
|
||||||
|
and have experimented in my own time with many different languages. In my
|
||||||
|
career I have used Python, .NET, PowerShell, TSQL, VB and SAS (multiple
|
||||||
|
products) in an Enterprise capacity. I also have experience using Google Cloud
|
||||||
|
Platform and AWS tools for ETL and data platform development as well as git
|
||||||
|
for version control and deployment using various IAC tools. I have also
|
||||||
|
conducted data analysis and modelling on business metrics to find relationships
|
||||||
|
between both staff and customer behavior and produced actionable
|
||||||
|
recommendations based on the conclusions. In a private context I have also
|
||||||
|
experimented with C, C# and Kotlin I am looking to further my career by taking
|
||||||
|
my passion for data engineering and analysis as well as web and software
|
||||||
|
development and applying it in a strategic context.</p>
|
||||||
|
<h1>SKILLS &amp; ABILITIES</h1>
|
||||||
<ul>
|
<ul>
|
||||||
<li>SAS</li>
|
<li>Python (scripting, compiling, notebooks – Sagemaker, Jupyter)</li>
|
||||||
<li>SSIS/SSRS</li>
|
<li>git</li>
|
||||||
<li>COGNOS/TM1</li>
|
<li>SAS (Base, EG, VA)</li>
|
||||||
|
<li>Various Google Cloud Tools (Data Fusion, Compute Engine, Cloud Functions)</li>
|
||||||
|
<li>Various Amazon Tools (EC2, RDS, Kinesis, Glue, Redshift, Lambda, ECS, ECR, EKS)</li>
|
||||||
|
<li>Streaming Technologies (Kafka, Hive, Spark Streaming)</li>
|
||||||
|
<li>Various DB platforms both on Prem and Serverless (MariaDB/MySql,</li>
|
||||||
|
<li>Postgres/Redshift, SQL Server, RDS/Aurora variants)</li>
|
||||||
|
<li>Various Microsoft Products (PowerBI, TSQL, Excel, VBA)</li>
|
||||||
|
<li>Linux Server Administration (cron, bash, systemD)</li>
|
||||||
|
<li>ETL/ELT Development</li>
|
||||||
|
<li>Basic Data Modelling (Kimball, SCD Type 2)</li>
|
||||||
|
<li>IAC (Cloud Formation, Terraform)</li>
|
||||||
|
<li>Datahub Deployment</li>
|
||||||
|
<li>Dagster Orchestration Deployments</li>
|
||||||
|
<li>DBT Modelling and Design Deployments</li>
|
||||||
|
<li>Containerised and Cloud Driven Data Architecture</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>Rickety:</em></p>
|
<h1>EXPERIENCE</h1>
|
||||||
|
<h2>Cloud Data Architect</h2>
|
||||||
|
<h3><em>Redeye Apps</em></h3>
|
||||||
|
<h4><em>May 2022 - Present</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Just write stored procedures!</li>
|
<li>Greenfields Research, Design and Deployment of S3 datalake (Parquet)</li>
|
||||||
<li>Startup script on my laptop XD</li>
|
<li>AWS DMS, S3, Athena, Glue</li>
|
||||||
<li>"Don't touch that machine over there, No one knows what it does but if it's turned off our financial reports don't work" (This is a third hand story I heard, seriously!)</li>
|
<li>Research Design and Deployment of Catalog (Datahub)</li>
|
||||||
<li>"I need to an upgrade to my laptop, Excel needs more than 8GB of RAM"</li>
|
<li>Design of Data Governance Process (Datahub driven)</li>
|
||||||
|
<li>Research Design and Deployment of Orchestration and Modelling for Transforms (Dagster/DBT into Mesos)</li>
|
||||||
|
<li>CI/CD design and deployment of modelling and orchestration using Gitlab</li>
|
||||||
|
<li>Research, Design and Deployment of ML Ops Dev pipelines anddeployment strategy</li>
|
||||||
|
<li>Design of ETL/Pipelines (DBT)</li>
|
||||||
|
<li>Design of Customer Facing Data Products and deployment methodologies (Fully automated via Kakfa/Dagster/DBT)</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>Hard:</em></p>
|
<h2>Data Engineer,</h2>
|
||||||
|
<h3><em>TechConnect IT Solutions</em></h3>
|
||||||
|
<h4><em>August 2021 – May 2022</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Hadoop</li>
|
<li>Design of Cloud Data Batch ETL solutions using Python (Glue)</li>
|
||||||
<li>Spark (hadoop but whatever)</li>
|
<li>Design of Cloud Data Streaming ETL solution using Python (Kinesis)</li>
|
||||||
<li>Python</li>
|
<li>Solve complex client business problems using software to join and transform data from DB’s, Web API’s, Application API’s and System logs</li>
|
||||||
<li>R</li>
|
<li>Build CI/CD pipelines to ensure smooth deployments (Bitbucket, gitlab)</li>
|
||||||
|
<li>Apply Prebuilt ML models to software solutions (Sagemaker)</li>
|
||||||
|
<li>Assist with the architecting of Containerisation solutions (Docker, ECS, ECR)</li>
|
||||||
|
<li>API testing and development (gRPC, Rest)</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>(The reason I've listed them as hard is because self hosting Hadoop/Spark and managing a truckload of python or R scripts, whilst it could have been "cheap" required a team of devs who really really knew what they were doing... so not really cheap and also <strong>really hard</strong>)</em></p>
|
<h2>Enterprise Data Warehouse Developer</h2>
|
||||||
<p>Then there was getting git behind all the sql scripts and modelling, let alone CI/CD <strong>IF</strong> it existed, it was custom, and bespoke and likely had a single point of failure in the person who knew how <code>git merge</code> worked. At least... thats I was told I'm not <em>that</em> old ;p.</p>
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
<p>These days we are pretty blessed, with democrotisation of clusters and data Infrastructure in the cloud we no longer need a team of sysadmins who know how to tune a cluster to the Nth degree to get the best our of our data workloads (well... we do, but we pay the cloud guys for that!). However, we still need to know about the idiosyncracities of this infrastructure, when it is appropiate to use and how we want to control and maintain the workloads. </p>
|
<h4><em>August 2019 - August 2021</em></h4>
|
||||||
<p>In general when I am designing a system I normally like to break it into 3.</p>
|
|
||||||
<ul>
|
<ul>
|
||||||
<li>Storage</li>
|
<li>ETL development of CRM, WFP, Outbound Dialer, Inbound switch in Google Cloud, SAS, TSQL</li>
|
||||||
<li>Compute</li>
|
<li>Bringing new data to the business to analyse for new insights</li>
|
||||||
<li>Code</li>
|
<li>Redeveloped Version Control and brought git to the data team</li>
|
||||||
|
<li>Introduced python for API enablement in the Enterprise Data Warehouse</li>
|
||||||
|
<li>Partnering with the business to focus data project on actual need and translating into technical requirements</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>In General</em> Storage and Compute will be infrastructer related, "Code" is sort of a catch all for my modelling, normally sql, python/r or spark scripts that are used to provide system or business logic, anything really thats going to get data to the end user/analyst/data scientists/annoying person </p>
|
<h2>Business Analyst</h2>
|
||||||
<p>Traditionally the compute layer only really had 2 considerations</p>
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2018 - August 2019</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>SQL or Logic engine (normally a flavour of spark(glue) and then something like reshift/athena/trino/bigquery)</li>
|
<li>Automate Service Performance Reporting using PowerShell/VBA/SAS</li>
|
||||||
<li>Orchestration Layer (Airflow, Dagster)</li>
|
<li>Learn and leverage SAS EG and VA to streamline Microsoft Excel Reporting</li>
|
||||||
|
<li>Identify and develop data pipelines to source data from multiple sources easily and collate into a single source to identify relationships and trends</li>
|
||||||
|
<li>Technologies used include VBA, PowerShell, SQL, Web API’s, SAS</li>
|
||||||
|
<li>Where SAS is inappropriate use VBA to automate processes in Microsoft Access and Excel</li>
|
||||||
|
<li>Gather Requirements to build meaningful reporting solutions</li>
|
||||||
|
<li>Provide meaningful analysis on business performance and provide relevant presentations and reports to senior stakeholders.</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p>But with the advent of sql engine agnostic Modelling we potentially now need to also consider</p>
|
<h2>Forecasting and Capacity Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2017 – January 2018</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Model Compilation</li>
|
<li>Develop the outbound forecasting model for the Auto and General sales call center by analysing the relationship between customer decisions and workload drivers</li>
|
||||||
|
<li>This includes the complete data pipeline for the model from identifying and sourcing data, building the reporting and analysing the data and associated drivers.</li>
|
||||||
|
<li>Forecast inbound workload requirements for the Auto and General sales call center using time series analysis</li>
|
||||||
|
<li>Learn and leverage the Aspect Workforce Management System to ensure efficiency of forecast generation</li>
|
||||||
|
<li>Learn and leverage the capabilities of SAS Enterprise Guide to improve accuracy</li>
|
||||||
|
<li>Liaise with people across the business to ensure meaningful, accurate analysis is provided to senior stakeholders</li>
|
||||||
|
<li>Analyse monthly, weekly and intraday requirements and ensure forecast is accurately predicting workload for breaks, meetings and Leave</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p>Now on the surface it seems counterintuitive to seperate the models from the logic layer but lets consider the following scenario</p>
|
<h2>Senior HR Performance Analyst</h2>
|
||||||
<blockquote>
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
<p>Redshift is costing to much and is getting slow, we want to try bigquery
|
<h4><em>June 2016 - January 2017</em></h4>
|
||||||
How much investment will it be to change over</p>
|
<ul>
|
||||||
</blockquote>
|
<li>Harmonise various systems to develop a unified workforce reporting and analysis framework with appropriate metrics</li>
|
||||||
<p>Now, If the entirety of your modelling is stored and deployed to big query direct this would not only involve the investment of spinning up the big query account and either connecting or migrating your data over. You would also need to consider how in the bloody hell you convert all your existing models and workflows over.</p>
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
<p>With something like DBT or sqlmesh you change your compilation target and it's done for you. It also means the Data Engineer now doesn't need to necessarily understand the esoteric nature of the target, at least for simple models (which, lets be real, most are).</p>
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
<p>BUT, now we have <em>a lot</em> of software and infrastructure a simplified common datastack will look something like the below (Assuming ELT, ETL is a bit different but more or less needs the same components)</p>
|
</ul>
|
||||||
<p><img src="http://localhost:8000/images/DataStackSimplified.png" width="600" height="295" /></p></content><category term="Data Engineering"></category><category term="data engineering"></category><category term="DBT"></category><category term="Terraform"></category><category term="IAC"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
<h2>Workforce Business Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>July 2015 – June 2016</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop and refine current workforce analysis techniques and databases</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Act as liaison between shared service providers and executives and facilitate communication during the implementation of a payroll leave audit</li>
|
||||||
|
<li>Gather reporting requirements from various business areas and produce ad-hoc and regular reports as required</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EDUCATION</h1>
|
||||||
|
<ul>
|
||||||
|
<li>2011 Bachelor of Business Management, University of Queensland</li>
|
||||||
|
<li>2008 Bachelor of Arts, University of Queensland</li>
|
||||||
|
</ul>
|
||||||
|
<h1>REFERENCES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Anthony Stiller Lead Developer, Data warehousing, Queensland Health</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0428 038 031</em></p>
|
||||||
|
<ul>
|
||||||
|
<li>Jaime Brian Head of Cloud Ninjas, TechConnect</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0422 012 17</em></p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>Metabase and DuckDB</title><link href="http://localhost:8000/metabase-duckdb.html" rel="alternate"></link><published>2023-11-15T20:00:00+10:00</published><updated>2023-11-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-11-15:/metabase-duckdb.html</id><summary type="html"><p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p></summary><content type="html"><p>Ahhhh <a href="https://duckdb.org/">DuckDB</a> if you're even partly floating around in the data space you've probably been hearing ALOT about it and it's <em>"Datawarehouse on your laptop"</em> mantra. However, the OTHER application that sometimes gets missed is <em>"SQLite for OLAP workloads"</em> and it was this concept that once I grasped it gave me a very interesting idea.... What if we could take the very pretty Aggregate Layer of our Data(warehouse/LakeHouse/Lake) and put that data right next to presentation layer of the lake, reducing network latency and... hopefully... have presentation reports running over very large workloads in the blink of an eye. It might even be fast enough that it could be deployed and embedded </p>
|
||||||
|
<p>However, for this to work we need some form of conatinerised reporting application.... lucky for us there is <a href="https://www.metabase.com/">Metabase</a> which is a fantastic little reporting application that has an open core. So this got me thinking... Can I put these two applications together and create a Reporting Layer with report embedding capabilities that is deployable in the cluster and has a admin UI accesible over a web page all whilst keeping the data locked to our network?</p>
|
||||||
|
<h3>The Beginnings of an Idea</h3>
|
||||||
|
<p>Ok so... Big first question. Can Duckdb and Metabase talk? Well... not quite. But first lets take a quick look at the architecture we'll be employing here </p>
|
||||||
|
<p><img alt="Duckdb Architecture" height="auto" width="100%" src="http://localhost:8000/images/metabase_duckdb.png"></p>
|
||||||
|
<p>But you'll notice this pretty glossed over line, "Connector", that right there is the clincher. So what is this "Connector"?. </p>
|
||||||
|
<p>To Deep dive into this would take a whole blog so to give you something to quickly wrap your head around its the glue that will make metabase be able to query your data source. The reality is its a jdbc driver compiled against metabase. </p>
|
||||||
|
<p>Thankfully Metabase point you to a <a href="https://github.com/AlexR2D2/metabase_duckdb_driver">community driver</a> for linking to duckdb ( hopefully it will be brought into metabase proper sooner rather than later ) </p>
|
||||||
|
<p>Now the release of this driver is still compiled against 0.8 of duckdb and 0.9 is the latest stable but hopefully the <a href="https://github.com/AlexR2D2/metabase_duckdb_driver/pull/19">PR</a> for this will land very soon giving a good quick way to link to the latest and greatest in duckdb from metabase</p>
|
||||||
|
<h3>But How do we get Data?</h3>
|
||||||
|
<p>Brilliant, using the recomended DockerFile we can load up a metabase container with the duckdb driver pre built</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">46.2</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">AlexR2D2</span><span class="o">/</span><span class="n">metabase_duckdb_driver</span><span class="o">/</span><span class="n">releases</span><span class="o">/</span><span class="n">download</span><span class="o">/</span><span class="mf">0.1</span><span class="o">.</span><span class="mi">6</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;java&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;-jar&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/metabase.jar&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>Great Now the big question. How do we get the data into the damn thing. Interestingly initially when I was designing this I had the thought of leveraging the in memory capabilities of duckdb and pulling in from the parquet on s3 directly as needed, after all the cluster is on AWS so the s3 API requests should be unbelievably fast anyway so why bother with a persistent database? </p>
|
||||||
|
<p>Now that we have the default credentials chain it is trivial to call parquet from s3</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">SELECT</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="k">FROM</span><span class="w"> </span><span class="n">read_parquet</span><span class="p">(</span><span class="s1">&#39;s3://&lt;bucket&gt;/&lt;file&gt;&#39;</span><span class="p">);</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>However, if you're reading direct off parquet all of a sudden you need to consider the partioning and I also found out that, if the parquet is being actively written to at the time of quering, duckdb has a hissyfit about metadata not matching the query. Needless to say duckdb and streaming parquet are not happy bed fellows (<em>and frankly were not desined to be so this is ok</em>). And the idea of trying to explain all this to the run of the mill reporting analyst whom it is my hope is a business sort of person not tech honestly gave me hives.. so I had to make it easier</p>
|
||||||
|
<p>The compromise occured to me... the curated layer is only built daily for reporting, and using that, I could create a duckdb file on disk that could be loaded into the metabase container itself.</p>
|
||||||
|
<p>With some very simple python as an operation in our orchestrator I had a job that would read direct from our curated parquet and create a duckdb file with it.. without giving away to much the job primarily consisted of this </p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">def</span> <span class="nf">duckdb_builder</span><span class="p">(</span><span class="n">table</span><span class="p">):</span>
|
||||||
|
<span class="n">conn</span> <span class="o">=</span> <span class="n">duckdb</span><span class="o">.</span><span class="n">connect</span><span class="p">(</span><span class="s2">&quot;curated_duckdb.duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;CALL load_aws_credentials(&#39;</span><span class="si">{</span><span class="n">aws_profile</span><span class="si">}</span><span class="s2">&#39;)&quot;</span><span class="p">)</span>
|
||||||
|
<span class="c1">#This removes a lot of weirdass ANSI in logs you DO NOT WANT</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">execute</span><span class="p">(</span><span class="s2">&quot;PRAGMA enable_progress_bar=false&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;Create </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> in duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">sql</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">&quot;CREATE OR REPLACE TABLE </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> AS SELECT * FROM read_parquet(&#39;s3://</span><span class="si">{</span><span class="n">curated_bucket</span><span class="si">}</span><span class="s2">/</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2">/*&#39;)&quot;</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="n">sql</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> Created&quot;</span><span class="p">)</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And then an upload to an s3 bucket</p>
|
||||||
|
<p>This of course necessated a cron job baked in to the metabase container itself to actually pull the duckdb in every morning. After some carefuly analysis of time (because I'm do lazy to implement message queues) I set up a s3 cp job that could be cronned direct from the container itself. This gives us a self updating metabase container pulling with a duckdb backend for client facing reporting right in the interface. AND because of the fact the duckdb is baked right into the container... there are NO associated s3 or dpu costs (merely the cost of running a relatively large container)</p>
|
||||||
|
<p>The final Dockerfile looks like this</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">47.6</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">mkdir</span><span class="w"> </span><span class="o">-</span><span class="n">p</span><span class="w"> </span><span class="o">/</span><span class="n">duckdb_data</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">entrypoint</span><span class="o">.</span><span class="n">sh</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">helper_scripts</span><span class="o">/</span><span class="n">download_duckdb</span><span class="o">.</span><span class="n">py</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">update</span><span class="w"> </span><span class="o">-</span><span class="n">y</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">upgrade</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">python3</span><span class="w"> </span><span class="n">python3</span><span class="o">-</span><span class="n">pip</span><span class="w"> </span><span class="n">cron</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">pip3</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">boto3</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span><span class="n">l</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">cat</span><span class="p">;</span><span class="w"> </span><span class="n">echo</span><span class="w"> </span><span class="s2">&quot;0 */6 * * * python3 /home/helper_scripts/download_duckdb.py&quot;</span><span class="p">;</span><span class="w"> </span><span class="p">}</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;bash&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/entrypoint.sh&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And there we have it... an in memory containerised reporting solution with blazing fast capability to aggregate and build reports based on curated data direct from the business.. fully automated and deployable via CI/CD, that provides data updates daily.</p>
|
||||||
|
<p>Now the embedded part.. which isn't built yet but I'll make sure to update you once we have/if we do because the architecture is very exciting for an embbdedded reporting workflow that is deployable via CI/CD processes to applications. As a little taster I'll point you to the <a href="https://www.metabase.com/learn/administration/git-based-workflow">metabase documentation</a>, the unfortunate thing about it is Metabase <em>have</em> hidden this behind the enterprise license.. but I can absolutely see why. If we get to implementing this I'll be sure to update you here on the learnings.</p>
|
||||||
|
<p>Until then....</p></content><category term="Business Intelligence"></category><category term="data engineering"></category><category term="Metabase"></category><category term="DuckDB"></category><category term="embedded"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
||||||
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
||||||
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
||||||
<h3>Datalake Extraction Layer</h3>
|
<h3>Datalake Extraction Layer</h3>
|
||||||
|
|||||||
+203
-38
@@ -1,54 +1,219 @@
|
|||||||
<?xml version="1.0" encoding="utf-8"?>
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/all.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-06-15T20:00:00+10:00</updated><entry><title>CI/CD in Data Engineering</title><link href="http://localhost:8000/CI/CD%20in%20Data%20and%20Data%20Infrastructure.html" rel="alternate"></link><published>2023-06-15T20:00:00+10:00</published><updated>2023-06-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-06-15:/CI/CD in Data and Data Infrastructure.html</id><summary type="html"><p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p></summary><content type="html"><p>Data Engineering has traditionally been considered the bastard step child of work that would have once been considered Administrative in the Tech world. Predominately we write SQL and then deploy that SQL onto one or more Databases. In fact a lot of the traditional methodologies around data almost assume this is the core of how an organistation is managing the majority of it's data. In the last couple of years though there has been a very steady move towards having the Data Engineering workload of SQL move towards Software Engineering techniques. With the popularity of tools like <a href="https://www.dbtlabs.com">DBT</a> and the latest newcommer on the block, <a href="https://www.sqlmesh.com">SQL-MESH</a> The oppportunity has started to arise where we can align our Data Engineering workloads with different environments and move much more efficiently towards a Continous Integration and Deployment methodology in our workflows. </p>
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/all.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2024-03-13T20:00:00+10:00</updated><entry><title>A Cover Letter</title><link href="http://localhost:8000/cover-letter.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/cover-letter.html</id><summary type="html"><p>A Summary of what I've done and Where I'd like to go for prospective Employers</p></summary><content type="html"><p>To whom it may concern</p>
|
||||||
<p>For the Data Engineering space the move to the cloud has been a breath of fresh air (Not so in some other IT disciplines). I am relatively young, so I don't 100% remember but my experience has taught me that there were 3 options here not so long ago</p>
|
<p>My name is Andrew Ridgway and I am a Data and Technology professional looking to embark on the next step in my career.</p>
|
||||||
<p><em>Expensive:</em></p>
|
<p>I have over 10 years’ experience in System and Data Architecture, Data Modelling and Orchestration, Business and Technical Analysis and System and Development Process Design. Most of this has been in developing Cloud architectures and workloads on AWS and GCP Including ML workloads using Sagemaker. </p>
|
||||||
|
<p>In my current role I have Proposed, Designed and built the data platform currently used by business. This includes internal and external data products as well as the infrastructure and modelling to support these. This role has seen me liaise with stakeholders of all levels of the business from Analysts in the Customer Experience team right up to C suite executives and preparing material for board members. I understand the complexity of communicating complex system design to different level stakeholders and the complexities of involved in communicating to both technical and less technical employees particularly in relation to data and ML technologies. </p>
|
||||||
|
<p>I have also worked as a technical consultant to many businesses and have assisted with the design and implementation of systems for a wide range of industries including financial services, mining and retail. I understand the complexities created by regulation in these environments and understand that this can sometimes necessitate the use of technologies and designs, including legacy systems and designs, I wouldn’t normally use. I also have a passion of designing systems that enable these organisations to realise the benefits of CI/CD on workloads they would not traditionally use this capability. In particular I took a very traditional legacy Data Warehousing team and implemented a solution that meant version control was no longer controlled by a daily copy and paste of folders with dates on major updates. My solution involved establishing guidelines of use of git version control so that this could happen automatically as people committed new code to the core code base. As I have moved into cloud architecture I have made sure to use best practice and ensure everything I build isn’t considered production ready until it is in IAC and deployed through a CI/CD pipeline.</p>
|
||||||
|
<p>In a personal capacity I am an avid tech and ML enthusiast. I have designed my own cluster including monitoring and deployment that runs several services that my family uses including chat and DNS and am in the process of designing a “set and forget” system that will allows me to have multi user tenancies on hardware I operate that should enable us to have the niceties of cloud services like email, storage and scheduling with the safety of knowing where that data is stored and exactly how it is used. I also like to design small IoT devices out of Arduino boards allowing me to monitor and control different facets of our house like temperature and light. </p>
|
||||||
|
<p>Currently I am working on a project to merge my skill in SQL Modelling and Orchestration with GPT API’s to try and lessen that burden. You can see some of this work in its very early stages here:
|
||||||
|
(gpt-sql-generator)[https://github.com/armistace/gpt-sql-generator]
|
||||||
|
(dbt_sources_generator)[https://github.com/armistace/datahub_dbt_sources_generator]</p>
|
||||||
|
<p>I look forward to hearing from you soon.</p>
|
||||||
|
<p>Sincerely,</p>
|
||||||
|
<hr>
|
||||||
|
<p>Andrew Ridgway</p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>A Resume</title><link href="http://localhost:8000/resume.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/resume.html</id><summary type="html"><p>A Summary of My work Experience</p></summary><content type="html"><h1>OVERVIEW</h1>
|
||||||
|
<p>I am a Senior Data Engineer looking to transition my skills to Data and Solution
|
||||||
|
Architecting as well as project management. I have spent the better part of the
|
||||||
|
last decade refining my abilities in taking business requirements and turning
|
||||||
|
those into actionable data engineering, analytics, and software projects with
|
||||||
|
trackable metrics. I believe in agnosticism when it comes to coding languages
|
||||||
|
and have experimented in my own time with many different languages. In my
|
||||||
|
career I have used Python, .NET, PowerShell, TSQL, VB and SAS (multiple
|
||||||
|
products) in an Enterprise capacity. I also have experience using Google Cloud
|
||||||
|
Platform and AWS tools for ETL and data platform development as well as git
|
||||||
|
for version control and deployment using various IAC tools. I have also
|
||||||
|
conducted data analysis and modelling on business metrics to find relationships
|
||||||
|
between both staff and customer behavior and produced actionable
|
||||||
|
recommendations based on the conclusions. In a private context I have also
|
||||||
|
experimented with C, C# and Kotlin I am looking to further my career by taking
|
||||||
|
my passion for data engineering and analysis as well as web and software
|
||||||
|
development and applying it in a strategic context.</p>
|
||||||
|
<h1>SKILLS &amp; ABILITIES</h1>
|
||||||
<ul>
|
<ul>
|
||||||
<li>SAS</li>
|
<li>Python (scripting, compiling, notebooks – Sagemaker, Jupyter)</li>
|
||||||
<li>SSIS/SSRS</li>
|
<li>git</li>
|
||||||
<li>COGNOS/TM1</li>
|
<li>SAS (Base, EG, VA)</li>
|
||||||
|
<li>Various Google Cloud Tools (Data Fusion, Compute Engine, Cloud Functions)</li>
|
||||||
|
<li>Various Amazon Tools (EC2, RDS, Kinesis, Glue, Redshift, Lambda, ECS, ECR, EKS)</li>
|
||||||
|
<li>Streaming Technologies (Kafka, Hive, Spark Streaming)</li>
|
||||||
|
<li>Various DB platforms both on Prem and Serverless (MariaDB/MySql,</li>
|
||||||
|
<li>Postgres/Redshift, SQL Server, RDS/Aurora variants)</li>
|
||||||
|
<li>Various Microsoft Products (PowerBI, TSQL, Excel, VBA)</li>
|
||||||
|
<li>Linux Server Administration (cron, bash, systemD)</li>
|
||||||
|
<li>ETL/ELT Development</li>
|
||||||
|
<li>Basic Data Modelling (Kimball, SCD Type 2)</li>
|
||||||
|
<li>IAC (Cloud Formation, Terraform)</li>
|
||||||
|
<li>Datahub Deployment</li>
|
||||||
|
<li>Dagster Orchestration Deployments</li>
|
||||||
|
<li>DBT Modelling and Design Deployments</li>
|
||||||
|
<li>Containerised and Cloud Driven Data Architecture</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>Rickety:</em></p>
|
<h1>EXPERIENCE</h1>
|
||||||
|
<h2>Cloud Data Architect</h2>
|
||||||
|
<h3><em>Redeye Apps</em></h3>
|
||||||
|
<h4><em>May 2022 - Present</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Just write stored procedures!</li>
|
<li>Greenfields Research, Design and Deployment of S3 datalake (Parquet)</li>
|
||||||
<li>Startup script on my laptop XD</li>
|
<li>AWS DMS, S3, Athena, Glue</li>
|
||||||
<li>"Don't touch that machine over there, No one knows what it does but if it's turned off our financial reports don't work" (This is a third hand story I heard, seriously!)</li>
|
<li>Research Design and Deployment of Catalog (Datahub)</li>
|
||||||
<li>"I need to an upgrade to my laptop, Excel needs more than 8GB of RAM"</li>
|
<li>Design of Data Governance Process (Datahub driven)</li>
|
||||||
|
<li>Research Design and Deployment of Orchestration and Modelling for Transforms (Dagster/DBT into Mesos)</li>
|
||||||
|
<li>CI/CD design and deployment of modelling and orchestration using Gitlab</li>
|
||||||
|
<li>Research, Design and Deployment of ML Ops Dev pipelines anddeployment strategy</li>
|
||||||
|
<li>Design of ETL/Pipelines (DBT)</li>
|
||||||
|
<li>Design of Customer Facing Data Products and deployment methodologies (Fully automated via Kakfa/Dagster/DBT)</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>Hard:</em></p>
|
<h2>Data Engineer,</h2>
|
||||||
|
<h3><em>TechConnect IT Solutions</em></h3>
|
||||||
|
<h4><em>August 2021 – May 2022</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Hadoop</li>
|
<li>Design of Cloud Data Batch ETL solutions using Python (Glue)</li>
|
||||||
<li>Spark (hadoop but whatever)</li>
|
<li>Design of Cloud Data Streaming ETL solution using Python (Kinesis)</li>
|
||||||
<li>Python</li>
|
<li>Solve complex client business problems using software to join and transform data from DB’s, Web API’s, Application API’s and System logs</li>
|
||||||
<li>R</li>
|
<li>Build CI/CD pipelines to ensure smooth deployments (Bitbucket, gitlab)</li>
|
||||||
|
<li>Apply Prebuilt ML models to software solutions (Sagemaker)</li>
|
||||||
|
<li>Assist with the architecting of Containerisation solutions (Docker, ECS, ECR)</li>
|
||||||
|
<li>API testing and development (gRPC, Rest)</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>(The reason I've listed them as hard is because self hosting Hadoop/Spark and managing a truckload of python or R scripts, whilst it could have been "cheap" required a team of devs who really really knew what they were doing... so not really cheap and also <strong>really hard</strong>)</em></p>
|
<h2>Enterprise Data Warehouse Developer</h2>
|
||||||
<p>Then there was getting git behind all the sql scripts and modelling, let alone CI/CD <strong>IF</strong> it existed, it was custom, and bespoke and likely had a single point of failure in the person who knew how <code>git merge</code> worked. At least... thats I was told I'm not <em>that</em> old ;p.</p>
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
<p>These days we are pretty blessed, with democrotisation of clusters and data Infrastructure in the cloud we no longer need a team of sysadmins who know how to tune a cluster to the Nth degree to get the best our of our data workloads (well... we do, but we pay the cloud guys for that!). However, we still need to know about the idiosyncracities of this infrastructure, when it is appropiate to use and how we want to control and maintain the workloads. </p>
|
<h4><em>August 2019 - August 2021</em></h4>
|
||||||
<p>In general when I am designing a system I normally like to break it into 3.</p>
|
|
||||||
<ul>
|
<ul>
|
||||||
<li>Storage</li>
|
<li>ETL development of CRM, WFP, Outbound Dialer, Inbound switch in Google Cloud, SAS, TSQL</li>
|
||||||
<li>Compute</li>
|
<li>Bringing new data to the business to analyse for new insights</li>
|
||||||
<li>Code</li>
|
<li>Redeveloped Version Control and brought git to the data team</li>
|
||||||
|
<li>Introduced python for API enablement in the Enterprise Data Warehouse</li>
|
||||||
|
<li>Partnering with the business to focus data project on actual need and translating into technical requirements</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>In General</em> Storage and Compute will be infrastructer related, "Code" is sort of a catch all for my modelling, normally sql, python/r or spark scripts that are used to provide system or business logic, anything really thats going to get data to the end user/analyst/data scientists/annoying person </p>
|
<h2>Business Analyst</h2>
|
||||||
<p>Traditionally the compute layer only really had 2 considerations</p>
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2018 - August 2019</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>SQL or Logic engine (normally a flavour of spark(glue) and then something like reshift/athena/trino/bigquery)</li>
|
<li>Automate Service Performance Reporting using PowerShell/VBA/SAS</li>
|
||||||
<li>Orchestration Layer (Airflow, Dagster)</li>
|
<li>Learn and leverage SAS EG and VA to streamline Microsoft Excel Reporting</li>
|
||||||
|
<li>Identify and develop data pipelines to source data from multiple sources easily and collate into a single source to identify relationships and trends</li>
|
||||||
|
<li>Technologies used include VBA, PowerShell, SQL, Web API’s, SAS</li>
|
||||||
|
<li>Where SAS is inappropriate use VBA to automate processes in Microsoft Access and Excel</li>
|
||||||
|
<li>Gather Requirements to build meaningful reporting solutions</li>
|
||||||
|
<li>Provide meaningful analysis on business performance and provide relevant presentations and reports to senior stakeholders.</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p>But with the advent of sql engine agnostic Modelling we potentially now need to also consider</p>
|
<h2>Forecasting and Capacity Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2017 – January 2018</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Model Compilation</li>
|
<li>Develop the outbound forecasting model for the Auto and General sales call center by analysing the relationship between customer decisions and workload drivers</li>
|
||||||
|
<li>This includes the complete data pipeline for the model from identifying and sourcing data, building the reporting and analysing the data and associated drivers.</li>
|
||||||
|
<li>Forecast inbound workload requirements for the Auto and General sales call center using time series analysis</li>
|
||||||
|
<li>Learn and leverage the Aspect Workforce Management System to ensure efficiency of forecast generation</li>
|
||||||
|
<li>Learn and leverage the capabilities of SAS Enterprise Guide to improve accuracy</li>
|
||||||
|
<li>Liaise with people across the business to ensure meaningful, accurate analysis is provided to senior stakeholders</li>
|
||||||
|
<li>Analyse monthly, weekly and intraday requirements and ensure forecast is accurately predicting workload for breaks, meetings and Leave</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p>Now on the surface it seems counterintuitive to seperate the models from the logic layer but lets consider the following scenario</p>
|
<h2>Senior HR Performance Analyst</h2>
|
||||||
<blockquote>
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
<p>Redshift is costing to much and is getting slow, we want to try bigquery
|
<h4><em>June 2016 - January 2017</em></h4>
|
||||||
How much investment will it be to change over</p>
|
<ul>
|
||||||
</blockquote>
|
<li>Harmonise various systems to develop a unified workforce reporting and analysis framework with appropriate metrics</li>
|
||||||
<p>Now, If the entirety of your modelling is stored and deployed to big query direct this would not only involve the investment of spinning up the big query account and either connecting or migrating your data over. You would also need to consider how in the bloody hell you convert all your existing models and workflows over.</p>
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
<p>With something like DBT or sqlmesh you change your compilation target and it's done for you. It also means the Data Engineer now doesn't need to necessarily understand the esoteric nature of the target, at least for simple models (which, lets be real, most are).</p>
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
<p>BUT, now we have <em>a lot</em> of software and infrastructure a simplified common datastack will look something like the below (Assuming ELT, ETL is a bit different but more or less needs the same components)</p>
|
</ul>
|
||||||
<p><img src="http://localhost:8000/images/DataStackSimplified.png" width="600" height="295" /></p></content><category term="Data Engineering"></category><category term="data engineering"></category><category term="DBT"></category><category term="Terraform"></category><category term="IAC"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
<h2>Workforce Business Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>July 2015 – June 2016</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop and refine current workforce analysis techniques and databases</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Act as liaison between shared service providers and executives and facilitate communication during the implementation of a payroll leave audit</li>
|
||||||
|
<li>Gather reporting requirements from various business areas and produce ad-hoc and regular reports as required</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EDUCATION</h1>
|
||||||
|
<ul>
|
||||||
|
<li>2011 Bachelor of Business Management, University of Queensland</li>
|
||||||
|
<li>2008 Bachelor of Arts, University of Queensland</li>
|
||||||
|
</ul>
|
||||||
|
<h1>REFERENCES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Anthony Stiller Lead Developer, Data warehousing, Queensland Health</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0428 038 031</em></p>
|
||||||
|
<ul>
|
||||||
|
<li>Jaime Brian Head of Cloud Ninjas, TechConnect</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0422 012 17</em></p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>Metabase and DuckDB</title><link href="http://localhost:8000/metabase-duckdb.html" rel="alternate"></link><published>2023-11-15T20:00:00+10:00</published><updated>2023-11-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-11-15:/metabase-duckdb.html</id><summary type="html"><p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p></summary><content type="html"><p>Ahhhh <a href="https://duckdb.org/">DuckDB</a> if you're even partly floating around in the data space you've probably been hearing ALOT about it and it's <em>"Datawarehouse on your laptop"</em> mantra. However, the OTHER application that sometimes gets missed is <em>"SQLite for OLAP workloads"</em> and it was this concept that once I grasped it gave me a very interesting idea.... What if we could take the very pretty Aggregate Layer of our Data(warehouse/LakeHouse/Lake) and put that data right next to presentation layer of the lake, reducing network latency and... hopefully... have presentation reports running over very large workloads in the blink of an eye. It might even be fast enough that it could be deployed and embedded </p>
|
||||||
|
<p>However, for this to work we need some form of conatinerised reporting application.... lucky for us there is <a href="https://www.metabase.com/">Metabase</a> which is a fantastic little reporting application that has an open core. So this got me thinking... Can I put these two applications together and create a Reporting Layer with report embedding capabilities that is deployable in the cluster and has a admin UI accesible over a web page all whilst keeping the data locked to our network?</p>
|
||||||
|
<h3>The Beginnings of an Idea</h3>
|
||||||
|
<p>Ok so... Big first question. Can Duckdb and Metabase talk? Well... not quite. But first lets take a quick look at the architecture we'll be employing here </p>
|
||||||
|
<p><img alt="Duckdb Architecture" height="auto" width="100%" src="http://localhost:8000/images/metabase_duckdb.png"></p>
|
||||||
|
<p>But you'll notice this pretty glossed over line, "Connector", that right there is the clincher. So what is this "Connector"?. </p>
|
||||||
|
<p>To Deep dive into this would take a whole blog so to give you something to quickly wrap your head around its the glue that will make metabase be able to query your data source. The reality is its a jdbc driver compiled against metabase. </p>
|
||||||
|
<p>Thankfully Metabase point you to a <a href="https://github.com/AlexR2D2/metabase_duckdb_driver">community driver</a> for linking to duckdb ( hopefully it will be brought into metabase proper sooner rather than later ) </p>
|
||||||
|
<p>Now the release of this driver is still compiled against 0.8 of duckdb and 0.9 is the latest stable but hopefully the <a href="https://github.com/AlexR2D2/metabase_duckdb_driver/pull/19">PR</a> for this will land very soon giving a good quick way to link to the latest and greatest in duckdb from metabase</p>
|
||||||
|
<h3>But How do we get Data?</h3>
|
||||||
|
<p>Brilliant, using the recomended DockerFile we can load up a metabase container with the duckdb driver pre built</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">46.2</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">AlexR2D2</span><span class="o">/</span><span class="n">metabase_duckdb_driver</span><span class="o">/</span><span class="n">releases</span><span class="o">/</span><span class="n">download</span><span class="o">/</span><span class="mf">0.1</span><span class="o">.</span><span class="mi">6</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;java&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;-jar&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/metabase.jar&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>Great Now the big question. How do we get the data into the damn thing. Interestingly initially when I was designing this I had the thought of leveraging the in memory capabilities of duckdb and pulling in from the parquet on s3 directly as needed, after all the cluster is on AWS so the s3 API requests should be unbelievably fast anyway so why bother with a persistent database? </p>
|
||||||
|
<p>Now that we have the default credentials chain it is trivial to call parquet from s3</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">SELECT</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="k">FROM</span><span class="w"> </span><span class="n">read_parquet</span><span class="p">(</span><span class="s1">&#39;s3://&lt;bucket&gt;/&lt;file&gt;&#39;</span><span class="p">);</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>However, if you're reading direct off parquet all of a sudden you need to consider the partioning and I also found out that, if the parquet is being actively written to at the time of quering, duckdb has a hissyfit about metadata not matching the query. Needless to say duckdb and streaming parquet are not happy bed fellows (<em>and frankly were not desined to be so this is ok</em>). And the idea of trying to explain all this to the run of the mill reporting analyst whom it is my hope is a business sort of person not tech honestly gave me hives.. so I had to make it easier</p>
|
||||||
|
<p>The compromise occured to me... the curated layer is only built daily for reporting, and using that, I could create a duckdb file on disk that could be loaded into the metabase container itself.</p>
|
||||||
|
<p>With some very simple python as an operation in our orchestrator I had a job that would read direct from our curated parquet and create a duckdb file with it.. without giving away to much the job primarily consisted of this </p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">def</span> <span class="nf">duckdb_builder</span><span class="p">(</span><span class="n">table</span><span class="p">):</span>
|
||||||
|
<span class="n">conn</span> <span class="o">=</span> <span class="n">duckdb</span><span class="o">.</span><span class="n">connect</span><span class="p">(</span><span class="s2">&quot;curated_duckdb.duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;CALL load_aws_credentials(&#39;</span><span class="si">{</span><span class="n">aws_profile</span><span class="si">}</span><span class="s2">&#39;)&quot;</span><span class="p">)</span>
|
||||||
|
<span class="c1">#This removes a lot of weirdass ANSI in logs you DO NOT WANT</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">execute</span><span class="p">(</span><span class="s2">&quot;PRAGMA enable_progress_bar=false&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;Create </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> in duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">sql</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">&quot;CREATE OR REPLACE TABLE </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> AS SELECT * FROM read_parquet(&#39;s3://</span><span class="si">{</span><span class="n">curated_bucket</span><span class="si">}</span><span class="s2">/</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2">/*&#39;)&quot;</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="n">sql</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> Created&quot;</span><span class="p">)</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And then an upload to an s3 bucket</p>
|
||||||
|
<p>This of course necessated a cron job baked in to the metabase container itself to actually pull the duckdb in every morning. After some carefuly analysis of time (because I'm do lazy to implement message queues) I set up a s3 cp job that could be cronned direct from the container itself. This gives us a self updating metabase container pulling with a duckdb backend for client facing reporting right in the interface. AND because of the fact the duckdb is baked right into the container... there are NO associated s3 or dpu costs (merely the cost of running a relatively large container)</p>
|
||||||
|
<p>The final Dockerfile looks like this</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">47.6</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">mkdir</span><span class="w"> </span><span class="o">-</span><span class="n">p</span><span class="w"> </span><span class="o">/</span><span class="n">duckdb_data</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">entrypoint</span><span class="o">.</span><span class="n">sh</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">helper_scripts</span><span class="o">/</span><span class="n">download_duckdb</span><span class="o">.</span><span class="n">py</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">update</span><span class="w"> </span><span class="o">-</span><span class="n">y</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">upgrade</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">python3</span><span class="w"> </span><span class="n">python3</span><span class="o">-</span><span class="n">pip</span><span class="w"> </span><span class="n">cron</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">pip3</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">boto3</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span><span class="n">l</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">cat</span><span class="p">;</span><span class="w"> </span><span class="n">echo</span><span class="w"> </span><span class="s2">&quot;0 */6 * * * python3 /home/helper_scripts/download_duckdb.py&quot;</span><span class="p">;</span><span class="w"> </span><span class="p">}</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;bash&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/entrypoint.sh&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And there we have it... an in memory containerised reporting solution with blazing fast capability to aggregate and build reports based on curated data direct from the business.. fully automated and deployable via CI/CD, that provides data updates daily.</p>
|
||||||
|
<p>Now the embedded part.. which isn't built yet but I'll make sure to update you once we have/if we do because the architecture is very exciting for an embbdedded reporting workflow that is deployable via CI/CD processes to applications. As a little taster I'll point you to the <a href="https://www.metabase.com/learn/administration/git-based-workflow">metabase documentation</a>, the unfortunate thing about it is Metabase <em>have</em> hidden this behind the enterprise license.. but I can absolutely see why. If we get to implementing this I'll be sure to update you here on the learnings.</p>
|
||||||
|
<p>Until then....</p></content><category term="Business Intelligence"></category><category term="data engineering"></category><category term="Metabase"></category><category term="DuckDB"></category><category term="embedded"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
||||||
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
||||||
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
||||||
<h3>Datalake Extraction Layer</h3>
|
<h3>Datalake Extraction Layer</h3>
|
||||||
|
|||||||
@@ -1,54 +1,219 @@
|
|||||||
<?xml version="1.0" encoding="utf-8"?>
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Andrew Ridgway</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/andrew-ridgway.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-06-15T20:00:00+10:00</updated><entry><title>CI/CD in Data Engineering</title><link href="http://localhost:8000/CI/CD%20in%20Data%20and%20Data%20Infrastructure.html" rel="alternate"></link><published>2023-06-15T20:00:00+10:00</published><updated>2023-06-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-06-15:/CI/CD in Data and Data Infrastructure.html</id><summary type="html"><p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p></summary><content type="html"><p>Data Engineering has traditionally been considered the bastard step child of work that would have once been considered Administrative in the Tech world. Predominately we write SQL and then deploy that SQL onto one or more Databases. In fact a lot of the traditional methodologies around data almost assume this is the core of how an organistation is managing the majority of it's data. In the last couple of years though there has been a very steady move towards having the Data Engineering workload of SQL move towards Software Engineering techniques. With the popularity of tools like <a href="https://www.dbtlabs.com">DBT</a> and the latest newcommer on the block, <a href="https://www.sqlmesh.com">SQL-MESH</a> The oppportunity has started to arise where we can align our Data Engineering workloads with different environments and move much more efficiently towards a Continous Integration and Deployment methodology in our workflows. </p>
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Andrew Ridgway</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/andrew-ridgway.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2024-03-13T20:00:00+10:00</updated><entry><title>A Cover Letter</title><link href="http://localhost:8000/cover-letter.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/cover-letter.html</id><summary type="html"><p>A Summary of what I've done and Where I'd like to go for prospective Employers</p></summary><content type="html"><p>To whom it may concern</p>
|
||||||
<p>For the Data Engineering space the move to the cloud has been a breath of fresh air (Not so in some other IT disciplines). I am relatively young, so I don't 100% remember but my experience has taught me that there were 3 options here not so long ago</p>
|
<p>My name is Andrew Ridgway and I am a Data and Technology professional looking to embark on the next step in my career.</p>
|
||||||
<p><em>Expensive:</em></p>
|
<p>I have over 10 years’ experience in System and Data Architecture, Data Modelling and Orchestration, Business and Technical Analysis and System and Development Process Design. Most of this has been in developing Cloud architectures and workloads on AWS and GCP Including ML workloads using Sagemaker. </p>
|
||||||
|
<p>In my current role I have Proposed, Designed and built the data platform currently used by business. This includes internal and external data products as well as the infrastructure and modelling to support these. This role has seen me liaise with stakeholders of all levels of the business from Analysts in the Customer Experience team right up to C suite executives and preparing material for board members. I understand the complexity of communicating complex system design to different level stakeholders and the complexities of involved in communicating to both technical and less technical employees particularly in relation to data and ML technologies. </p>
|
||||||
|
<p>I have also worked as a technical consultant to many businesses and have assisted with the design and implementation of systems for a wide range of industries including financial services, mining and retail. I understand the complexities created by regulation in these environments and understand that this can sometimes necessitate the use of technologies and designs, including legacy systems and designs, I wouldn’t normally use. I also have a passion of designing systems that enable these organisations to realise the benefits of CI/CD on workloads they would not traditionally use this capability. In particular I took a very traditional legacy Data Warehousing team and implemented a solution that meant version control was no longer controlled by a daily copy and paste of folders with dates on major updates. My solution involved establishing guidelines of use of git version control so that this could happen automatically as people committed new code to the core code base. As I have moved into cloud architecture I have made sure to use best practice and ensure everything I build isn’t considered production ready until it is in IAC and deployed through a CI/CD pipeline.</p>
|
||||||
|
<p>In a personal capacity I am an avid tech and ML enthusiast. I have designed my own cluster including monitoring and deployment that runs several services that my family uses including chat and DNS and am in the process of designing a “set and forget” system that will allows me to have multi user tenancies on hardware I operate that should enable us to have the niceties of cloud services like email, storage and scheduling with the safety of knowing where that data is stored and exactly how it is used. I also like to design small IoT devices out of Arduino boards allowing me to monitor and control different facets of our house like temperature and light. </p>
|
||||||
|
<p>Currently I am working on a project to merge my skill in SQL Modelling and Orchestration with GPT API’s to try and lessen that burden. You can see some of this work in its very early stages here:
|
||||||
|
(gpt-sql-generator)[https://github.com/armistace/gpt-sql-generator]
|
||||||
|
(dbt_sources_generator)[https://github.com/armistace/datahub_dbt_sources_generator]</p>
|
||||||
|
<p>I look forward to hearing from you soon.</p>
|
||||||
|
<p>Sincerely,</p>
|
||||||
|
<hr>
|
||||||
|
<p>Andrew Ridgway</p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>A Resume</title><link href="http://localhost:8000/resume.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/resume.html</id><summary type="html"><p>A Summary of My work Experience</p></summary><content type="html"><h1>OVERVIEW</h1>
|
||||||
|
<p>I am a Senior Data Engineer looking to transition my skills to Data and Solution
|
||||||
|
Architecting as well as project management. I have spent the better part of the
|
||||||
|
last decade refining my abilities in taking business requirements and turning
|
||||||
|
those into actionable data engineering, analytics, and software projects with
|
||||||
|
trackable metrics. I believe in agnosticism when it comes to coding languages
|
||||||
|
and have experimented in my own time with many different languages. In my
|
||||||
|
career I have used Python, .NET, PowerShell, TSQL, VB and SAS (multiple
|
||||||
|
products) in an Enterprise capacity. I also have experience using Google Cloud
|
||||||
|
Platform and AWS tools for ETL and data platform development as well as git
|
||||||
|
for version control and deployment using various IAC tools. I have also
|
||||||
|
conducted data analysis and modelling on business metrics to find relationships
|
||||||
|
between both staff and customer behavior and produced actionable
|
||||||
|
recommendations based on the conclusions. In a private context I have also
|
||||||
|
experimented with C, C# and Kotlin I am looking to further my career by taking
|
||||||
|
my passion for data engineering and analysis as well as web and software
|
||||||
|
development and applying it in a strategic context.</p>
|
||||||
|
<h1>SKILLS &amp; ABILITIES</h1>
|
||||||
<ul>
|
<ul>
|
||||||
<li>SAS</li>
|
<li>Python (scripting, compiling, notebooks – Sagemaker, Jupyter)</li>
|
||||||
<li>SSIS/SSRS</li>
|
<li>git</li>
|
||||||
<li>COGNOS/TM1</li>
|
<li>SAS (Base, EG, VA)</li>
|
||||||
|
<li>Various Google Cloud Tools (Data Fusion, Compute Engine, Cloud Functions)</li>
|
||||||
|
<li>Various Amazon Tools (EC2, RDS, Kinesis, Glue, Redshift, Lambda, ECS, ECR, EKS)</li>
|
||||||
|
<li>Streaming Technologies (Kafka, Hive, Spark Streaming)</li>
|
||||||
|
<li>Various DB platforms both on Prem and Serverless (MariaDB/MySql,</li>
|
||||||
|
<li>Postgres/Redshift, SQL Server, RDS/Aurora variants)</li>
|
||||||
|
<li>Various Microsoft Products (PowerBI, TSQL, Excel, VBA)</li>
|
||||||
|
<li>Linux Server Administration (cron, bash, systemD)</li>
|
||||||
|
<li>ETL/ELT Development</li>
|
||||||
|
<li>Basic Data Modelling (Kimball, SCD Type 2)</li>
|
||||||
|
<li>IAC (Cloud Formation, Terraform)</li>
|
||||||
|
<li>Datahub Deployment</li>
|
||||||
|
<li>Dagster Orchestration Deployments</li>
|
||||||
|
<li>DBT Modelling and Design Deployments</li>
|
||||||
|
<li>Containerised and Cloud Driven Data Architecture</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>Rickety:</em></p>
|
<h1>EXPERIENCE</h1>
|
||||||
|
<h2>Cloud Data Architect</h2>
|
||||||
|
<h3><em>Redeye Apps</em></h3>
|
||||||
|
<h4><em>May 2022 - Present</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Just write stored procedures!</li>
|
<li>Greenfields Research, Design and Deployment of S3 datalake (Parquet)</li>
|
||||||
<li>Startup script on my laptop XD</li>
|
<li>AWS DMS, S3, Athena, Glue</li>
|
||||||
<li>"Don't touch that machine over there, No one knows what it does but if it's turned off our financial reports don't work" (This is a third hand story I heard, seriously!)</li>
|
<li>Research Design and Deployment of Catalog (Datahub)</li>
|
||||||
<li>"I need to an upgrade to my laptop, Excel needs more than 8GB of RAM"</li>
|
<li>Design of Data Governance Process (Datahub driven)</li>
|
||||||
|
<li>Research Design and Deployment of Orchestration and Modelling for Transforms (Dagster/DBT into Mesos)</li>
|
||||||
|
<li>CI/CD design and deployment of modelling and orchestration using Gitlab</li>
|
||||||
|
<li>Research, Design and Deployment of ML Ops Dev pipelines anddeployment strategy</li>
|
||||||
|
<li>Design of ETL/Pipelines (DBT)</li>
|
||||||
|
<li>Design of Customer Facing Data Products and deployment methodologies (Fully automated via Kakfa/Dagster/DBT)</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>Hard:</em></p>
|
<h2>Data Engineer,</h2>
|
||||||
|
<h3><em>TechConnect IT Solutions</em></h3>
|
||||||
|
<h4><em>August 2021 – May 2022</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Hadoop</li>
|
<li>Design of Cloud Data Batch ETL solutions using Python (Glue)</li>
|
||||||
<li>Spark (hadoop but whatever)</li>
|
<li>Design of Cloud Data Streaming ETL solution using Python (Kinesis)</li>
|
||||||
<li>Python</li>
|
<li>Solve complex client business problems using software to join and transform data from DB’s, Web API’s, Application API’s and System logs</li>
|
||||||
<li>R</li>
|
<li>Build CI/CD pipelines to ensure smooth deployments (Bitbucket, gitlab)</li>
|
||||||
|
<li>Apply Prebuilt ML models to software solutions (Sagemaker)</li>
|
||||||
|
<li>Assist with the architecting of Containerisation solutions (Docker, ECS, ECR)</li>
|
||||||
|
<li>API testing and development (gRPC, Rest)</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>(The reason I've listed them as hard is because self hosting Hadoop/Spark and managing a truckload of python or R scripts, whilst it could have been "cheap" required a team of devs who really really knew what they were doing... so not really cheap and also <strong>really hard</strong>)</em></p>
|
<h2>Enterprise Data Warehouse Developer</h2>
|
||||||
<p>Then there was getting git behind all the sql scripts and modelling, let alone CI/CD <strong>IF</strong> it existed, it was custom, and bespoke and likely had a single point of failure in the person who knew how <code>git merge</code> worked. At least... thats I was told I'm not <em>that</em> old ;p.</p>
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
<p>These days we are pretty blessed, with democrotisation of clusters and data Infrastructure in the cloud we no longer need a team of sysadmins who know how to tune a cluster to the Nth degree to get the best our of our data workloads (well... we do, but we pay the cloud guys for that!). However, we still need to know about the idiosyncracities of this infrastructure, when it is appropiate to use and how we want to control and maintain the workloads. </p>
|
<h4><em>August 2019 - August 2021</em></h4>
|
||||||
<p>In general when I am designing a system I normally like to break it into 3.</p>
|
|
||||||
<ul>
|
<ul>
|
||||||
<li>Storage</li>
|
<li>ETL development of CRM, WFP, Outbound Dialer, Inbound switch in Google Cloud, SAS, TSQL</li>
|
||||||
<li>Compute</li>
|
<li>Bringing new data to the business to analyse for new insights</li>
|
||||||
<li>Code</li>
|
<li>Redeveloped Version Control and brought git to the data team</li>
|
||||||
|
<li>Introduced python for API enablement in the Enterprise Data Warehouse</li>
|
||||||
|
<li>Partnering with the business to focus data project on actual need and translating into technical requirements</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p><em>In General</em> Storage and Compute will be infrastructer related, "Code" is sort of a catch all for my modelling, normally sql, python/r or spark scripts that are used to provide system or business logic, anything really thats going to get data to the end user/analyst/data scientists/annoying person </p>
|
<h2>Business Analyst</h2>
|
||||||
<p>Traditionally the compute layer only really had 2 considerations</p>
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2018 - August 2019</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>SQL or Logic engine (normally a flavour of spark(glue) and then something like reshift/athena/trino/bigquery)</li>
|
<li>Automate Service Performance Reporting using PowerShell/VBA/SAS</li>
|
||||||
<li>Orchestration Layer (Airflow, Dagster)</li>
|
<li>Learn and leverage SAS EG and VA to streamline Microsoft Excel Reporting</li>
|
||||||
|
<li>Identify and develop data pipelines to source data from multiple sources easily and collate into a single source to identify relationships and trends</li>
|
||||||
|
<li>Technologies used include VBA, PowerShell, SQL, Web API’s, SAS</li>
|
||||||
|
<li>Where SAS is inappropriate use VBA to automate processes in Microsoft Access and Excel</li>
|
||||||
|
<li>Gather Requirements to build meaningful reporting solutions</li>
|
||||||
|
<li>Provide meaningful analysis on business performance and provide relevant presentations and reports to senior stakeholders.</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p>But with the advent of sql engine agnostic Modelling we potentially now need to also consider</p>
|
<h2>Forecasting and Capacity Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2017 – January 2018</em></h4>
|
||||||
<ul>
|
<ul>
|
||||||
<li>Model Compilation</li>
|
<li>Develop the outbound forecasting model for the Auto and General sales call center by analysing the relationship between customer decisions and workload drivers</li>
|
||||||
|
<li>This includes the complete data pipeline for the model from identifying and sourcing data, building the reporting and analysing the data and associated drivers.</li>
|
||||||
|
<li>Forecast inbound workload requirements for the Auto and General sales call center using time series analysis</li>
|
||||||
|
<li>Learn and leverage the Aspect Workforce Management System to ensure efficiency of forecast generation</li>
|
||||||
|
<li>Learn and leverage the capabilities of SAS Enterprise Guide to improve accuracy</li>
|
||||||
|
<li>Liaise with people across the business to ensure meaningful, accurate analysis is provided to senior stakeholders</li>
|
||||||
|
<li>Analyse monthly, weekly and intraday requirements and ensure forecast is accurately predicting workload for breaks, meetings and Leave</li>
|
||||||
</ul>
|
</ul>
|
||||||
<p>Now on the surface it seems counterintuitive to seperate the models from the logic layer but lets consider the following scenario</p>
|
<h2>Senior HR Performance Analyst</h2>
|
||||||
<blockquote>
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
<p>Redshift is costing to much and is getting slow, we want to try bigquery
|
<h4><em>June 2016 - January 2017</em></h4>
|
||||||
How much investment will it be to change over</p>
|
<ul>
|
||||||
</blockquote>
|
<li>Harmonise various systems to develop a unified workforce reporting and analysis framework with appropriate metrics</li>
|
||||||
<p>Now, If the entirety of your modelling is stored and deployed to big query direct this would not only involve the investment of spinning up the big query account and either connecting or migrating your data over. You would also need to consider how in the bloody hell you convert all your existing models and workflows over.</p>
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
<p>With something like DBT or sqlmesh you change your compilation target and it's done for you. It also means the Data Engineer now doesn't need to necessarily understand the esoteric nature of the target, at least for simple models (which, lets be real, most are).</p>
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
<p>BUT, now we have <em>a lot</em> of software and infrastructure a simplified common datastack will look something like the below (Assuming ELT, ETL is a bit different but more or less needs the same components)</p>
|
</ul>
|
||||||
<p><img src="http://localhost:8000/images/DataStackSimplified.png" width="600" height="295" /></p></content><category term="Data Engineering"></category><category term="data engineering"></category><category term="DBT"></category><category term="Terraform"></category><category term="IAC"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
<h2>Workforce Business Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>July 2015 – June 2016</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop and refine current workforce analysis techniques and databases</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Act as liaison between shared service providers and executives and facilitate communication during the implementation of a payroll leave audit</li>
|
||||||
|
<li>Gather reporting requirements from various business areas and produce ad-hoc and regular reports as required</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EDUCATION</h1>
|
||||||
|
<ul>
|
||||||
|
<li>2011 Bachelor of Business Management, University of Queensland</li>
|
||||||
|
<li>2008 Bachelor of Arts, University of Queensland</li>
|
||||||
|
</ul>
|
||||||
|
<h1>REFERENCES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Anthony Stiller Lead Developer, Data warehousing, Queensland Health</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0428 038 031</em></p>
|
||||||
|
<ul>
|
||||||
|
<li>Jaime Brian Head of Cloud Ninjas, TechConnect</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0422 012 17</em></p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>Metabase and DuckDB</title><link href="http://localhost:8000/metabase-duckdb.html" rel="alternate"></link><published>2023-11-15T20:00:00+10:00</published><updated>2023-11-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-11-15:/metabase-duckdb.html</id><summary type="html"><p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p></summary><content type="html"><p>Ahhhh <a href="https://duckdb.org/">DuckDB</a> if you're even partly floating around in the data space you've probably been hearing ALOT about it and it's <em>"Datawarehouse on your laptop"</em> mantra. However, the OTHER application that sometimes gets missed is <em>"SQLite for OLAP workloads"</em> and it was this concept that once I grasped it gave me a very interesting idea.... What if we could take the very pretty Aggregate Layer of our Data(warehouse/LakeHouse/Lake) and put that data right next to presentation layer of the lake, reducing network latency and... hopefully... have presentation reports running over very large workloads in the blink of an eye. It might even be fast enough that it could be deployed and embedded </p>
|
||||||
|
<p>However, for this to work we need some form of conatinerised reporting application.... lucky for us there is <a href="https://www.metabase.com/">Metabase</a> which is a fantastic little reporting application that has an open core. So this got me thinking... Can I put these two applications together and create a Reporting Layer with report embedding capabilities that is deployable in the cluster and has a admin UI accesible over a web page all whilst keeping the data locked to our network?</p>
|
||||||
|
<h3>The Beginnings of an Idea</h3>
|
||||||
|
<p>Ok so... Big first question. Can Duckdb and Metabase talk? Well... not quite. But first lets take a quick look at the architecture we'll be employing here </p>
|
||||||
|
<p><img alt="Duckdb Architecture" height="auto" width="100%" src="http://localhost:8000/images/metabase_duckdb.png"></p>
|
||||||
|
<p>But you'll notice this pretty glossed over line, "Connector", that right there is the clincher. So what is this "Connector"?. </p>
|
||||||
|
<p>To Deep dive into this would take a whole blog so to give you something to quickly wrap your head around its the glue that will make metabase be able to query your data source. The reality is its a jdbc driver compiled against metabase. </p>
|
||||||
|
<p>Thankfully Metabase point you to a <a href="https://github.com/AlexR2D2/metabase_duckdb_driver">community driver</a> for linking to duckdb ( hopefully it will be brought into metabase proper sooner rather than later ) </p>
|
||||||
|
<p>Now the release of this driver is still compiled against 0.8 of duckdb and 0.9 is the latest stable but hopefully the <a href="https://github.com/AlexR2D2/metabase_duckdb_driver/pull/19">PR</a> for this will land very soon giving a good quick way to link to the latest and greatest in duckdb from metabase</p>
|
||||||
|
<h3>But How do we get Data?</h3>
|
||||||
|
<p>Brilliant, using the recomended DockerFile we can load up a metabase container with the duckdb driver pre built</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">46.2</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">AlexR2D2</span><span class="o">/</span><span class="n">metabase_duckdb_driver</span><span class="o">/</span><span class="n">releases</span><span class="o">/</span><span class="n">download</span><span class="o">/</span><span class="mf">0.1</span><span class="o">.</span><span class="mi">6</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;java&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;-jar&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/metabase.jar&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>Great Now the big question. How do we get the data into the damn thing. Interestingly initially when I was designing this I had the thought of leveraging the in memory capabilities of duckdb and pulling in from the parquet on s3 directly as needed, after all the cluster is on AWS so the s3 API requests should be unbelievably fast anyway so why bother with a persistent database? </p>
|
||||||
|
<p>Now that we have the default credentials chain it is trivial to call parquet from s3</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">SELECT</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="k">FROM</span><span class="w"> </span><span class="n">read_parquet</span><span class="p">(</span><span class="s1">&#39;s3://&lt;bucket&gt;/&lt;file&gt;&#39;</span><span class="p">);</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>However, if you're reading direct off parquet all of a sudden you need to consider the partioning and I also found out that, if the parquet is being actively written to at the time of quering, duckdb has a hissyfit about metadata not matching the query. Needless to say duckdb and streaming parquet are not happy bed fellows (<em>and frankly were not desined to be so this is ok</em>). And the idea of trying to explain all this to the run of the mill reporting analyst whom it is my hope is a business sort of person not tech honestly gave me hives.. so I had to make it easier</p>
|
||||||
|
<p>The compromise occured to me... the curated layer is only built daily for reporting, and using that, I could create a duckdb file on disk that could be loaded into the metabase container itself.</p>
|
||||||
|
<p>With some very simple python as an operation in our orchestrator I had a job that would read direct from our curated parquet and create a duckdb file with it.. without giving away to much the job primarily consisted of this </p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">def</span> <span class="nf">duckdb_builder</span><span class="p">(</span><span class="n">table</span><span class="p">):</span>
|
||||||
|
<span class="n">conn</span> <span class="o">=</span> <span class="n">duckdb</span><span class="o">.</span><span class="n">connect</span><span class="p">(</span><span class="s2">&quot;curated_duckdb.duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;CALL load_aws_credentials(&#39;</span><span class="si">{</span><span class="n">aws_profile</span><span class="si">}</span><span class="s2">&#39;)&quot;</span><span class="p">)</span>
|
||||||
|
<span class="c1">#This removes a lot of weirdass ANSI in logs you DO NOT WANT</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">execute</span><span class="p">(</span><span class="s2">&quot;PRAGMA enable_progress_bar=false&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;Create </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> in duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">sql</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">&quot;CREATE OR REPLACE TABLE </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> AS SELECT * FROM read_parquet(&#39;s3://</span><span class="si">{</span><span class="n">curated_bucket</span><span class="si">}</span><span class="s2">/</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2">/*&#39;)&quot;</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="n">sql</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> Created&quot;</span><span class="p">)</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And then an upload to an s3 bucket</p>
|
||||||
|
<p>This of course necessated a cron job baked in to the metabase container itself to actually pull the duckdb in every morning. After some carefuly analysis of time (because I'm do lazy to implement message queues) I set up a s3 cp job that could be cronned direct from the container itself. This gives us a self updating metabase container pulling with a duckdb backend for client facing reporting right in the interface. AND because of the fact the duckdb is baked right into the container... there are NO associated s3 or dpu costs (merely the cost of running a relatively large container)</p>
|
||||||
|
<p>The final Dockerfile looks like this</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">47.6</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">mkdir</span><span class="w"> </span><span class="o">-</span><span class="n">p</span><span class="w"> </span><span class="o">/</span><span class="n">duckdb_data</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">entrypoint</span><span class="o">.</span><span class="n">sh</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">helper_scripts</span><span class="o">/</span><span class="n">download_duckdb</span><span class="o">.</span><span class="n">py</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">update</span><span class="w"> </span><span class="o">-</span><span class="n">y</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">upgrade</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">python3</span><span class="w"> </span><span class="n">python3</span><span class="o">-</span><span class="n">pip</span><span class="w"> </span><span class="n">cron</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">pip3</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">boto3</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span><span class="n">l</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">cat</span><span class="p">;</span><span class="w"> </span><span class="n">echo</span><span class="w"> </span><span class="s2">&quot;0 */6 * * * python3 /home/helper_scripts/download_duckdb.py&quot;</span><span class="p">;</span><span class="w"> </span><span class="p">}</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;bash&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/entrypoint.sh&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And there we have it... an in memory containerised reporting solution with blazing fast capability to aggregate and build reports based on curated data direct from the business.. fully automated and deployable via CI/CD, that provides data updates daily.</p>
|
||||||
|
<p>Now the embedded part.. which isn't built yet but I'll make sure to update you once we have/if we do because the architecture is very exciting for an embbdedded reporting workflow that is deployable via CI/CD processes to applications. As a little taster I'll point you to the <a href="https://www.metabase.com/learn/administration/git-based-workflow">metabase documentation</a>, the unfortunate thing about it is Metabase <em>have</em> hidden this behind the enterprise license.. but I can absolutely see why. If we get to implementing this I'll be sure to update you here on the learnings.</p>
|
||||||
|
<p>Until then....</p></content><category term="Business Intelligence"></category><category term="data engineering"></category><category term="Metabase"></category><category term="DuckDB"></category><category term="embedded"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
||||||
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
||||||
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
||||||
<h3>Datalake Extraction Layer</h3>
|
<h3>Datalake Extraction Layer</h3>
|
||||||
|
|||||||
@@ -1,2 +1,2 @@
|
|||||||
<?xml version="1.0" encoding="utf-8"?>
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
<rss version="2.0"><channel><title>Andrew Ridgway's Blog - Andrew Ridgway</title><link>http://localhost:8000/</link><description></description><lastBuildDate>Thu, 15 Jun 2023 20:00:00 +1000</lastBuildDate><item><title>CI/CD in Data Engineering</title><link>http://localhost:8000/CI/CD%20in%20Data%20and%20Data%20Infrastructure.html</link><description><p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Thu, 15 Jun 2023 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2023-06-15:/CI/CD in Data and Data Infrastructure.html</guid><category>Data Engineering</category><category>data engineering</category><category>DBT</category><category>Terraform</category><category>IAC</category></item><item><title>Implmenting Appflow in a Production Datalake</title><link>http://localhost:8000/appflow-production.html</link><description><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Tue, 23 May 2023 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2023-05-23:/appflow-production.html</guid><category>Data Engineering</category><category>data engineering</category><category>Amazon</category><category>Managed Services</category></item><item><title>Dawn of another blog attempt</title><link>http://localhost:8000/how-i-built-the-damn-thing.html</link><description><p>Containers and How I take my learnings from home and apply them to work</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Wed, 10 May 2023 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2023-05-10:/how-i-built-the-damn-thing.html</guid><category>Data Engineering</category><category>data engineering</category><category>containers</category></item></channel></rss>
|
<rss version="2.0"><channel><title>Andrew Ridgway's Blog - Andrew Ridgway</title><link>http://localhost:8000/</link><description></description><lastBuildDate>Wed, 13 Mar 2024 20:00:00 +1000</lastBuildDate><item><title>A Cover Letter</title><link>http://localhost:8000/cover-letter.html</link><description><p>A Summary of what I've done and Where I'd like to go for prospective Employers</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Fri, 23 Feb 2024 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2024-02-23:/cover-letter.html</guid><category>Resume</category><category>Cover Letter</category><category>Resume</category></item><item><title>A Resume</title><link>http://localhost:8000/resume.html</link><description><p>A Summary of My work Experience</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Fri, 23 Feb 2024 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2024-02-23:/resume.html</guid><category>Resume</category><category>Cover Letter</category><category>Resume</category></item><item><title>Metabase and DuckDB</title><link>http://localhost:8000/metabase-duckdb.html</link><description><p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Wed, 15 Nov 2023 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2023-11-15:/metabase-duckdb.html</guid><category>Business Intelligence</category><category>data engineering</category><category>Metabase</category><category>DuckDB</category><category>embedded</category></item><item><title>Implmenting Appflow in a Production Datalake</title><link>http://localhost:8000/appflow-production.html</link><description><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Tue, 23 May 2023 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2023-05-23:/appflow-production.html</guid><category>Data Engineering</category><category>data engineering</category><category>Amazon</category><category>Managed Services</category></item><item><title>Dawn of another blog attempt</title><link>http://localhost:8000/how-i-built-the-damn-thing.html</link><description><p>Containers and How I take my learnings from home and apply them to work</p></description><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Andrew Ridgway</dc:creator><pubDate>Wed, 10 May 2023 20:00:00 +1000</pubDate><guid isPermaLink="false">tag:localhost,2023-05-10:/how-i-built-the-damn-thing.html</guid><category>Data Engineering</category><category>data engineering</category><category>containers</category></item></channel></rss>
|
||||||
@@ -0,0 +1,75 @@
|
|||||||
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Business Intelligence</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/business-intelligence.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-11-15T20:00:00+10:00</updated><entry><title>Metabase and DuckDB</title><link href="http://localhost:8000/metabase-duckdb.html" rel="alternate"></link><published>2023-11-15T20:00:00+10:00</published><updated>2023-11-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-11-15:/metabase-duckdb.html</id><summary type="html"><p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p></summary><content type="html"><p>Ahhhh <a href="https://duckdb.org/">DuckDB</a> if you're even partly floating around in the data space you've probably been hearing ALOT about it and it's <em>"Datawarehouse on your laptop"</em> mantra. However, the OTHER application that sometimes gets missed is <em>"SQLite for OLAP workloads"</em> and it was this concept that once I grasped it gave me a very interesting idea.... What if we could take the very pretty Aggregate Layer of our Data(warehouse/LakeHouse/Lake) and put that data right next to presentation layer of the lake, reducing network latency and... hopefully... have presentation reports running over very large workloads in the blink of an eye. It might even be fast enough that it could be deployed and embedded </p>
|
||||||
|
<p>However, for this to work we need some form of conatinerised reporting application.... lucky for us there is <a href="https://www.metabase.com/">Metabase</a> which is a fantastic little reporting application that has an open core. So this got me thinking... Can I put these two applications together and create a Reporting Layer with report embedding capabilities that is deployable in the cluster and has a admin UI accesible over a web page all whilst keeping the data locked to our network?</p>
|
||||||
|
<h3>The Beginnings of an Idea</h3>
|
||||||
|
<p>Ok so... Big first question. Can Duckdb and Metabase talk? Well... not quite. But first lets take a quick look at the architecture we'll be employing here </p>
|
||||||
|
<p><img alt="Duckdb Architecture" height="auto" width="100%" src="http://localhost:8000/images/metabase_duckdb.png"></p>
|
||||||
|
<p>But you'll notice this pretty glossed over line, "Connector", that right there is the clincher. So what is this "Connector"?. </p>
|
||||||
|
<p>To Deep dive into this would take a whole blog so to give you something to quickly wrap your head around its the glue that will make metabase be able to query your data source. The reality is its a jdbc driver compiled against metabase. </p>
|
||||||
|
<p>Thankfully Metabase point you to a <a href="https://github.com/AlexR2D2/metabase_duckdb_driver">community driver</a> for linking to duckdb ( hopefully it will be brought into metabase proper sooner rather than later ) </p>
|
||||||
|
<p>Now the release of this driver is still compiled against 0.8 of duckdb and 0.9 is the latest stable but hopefully the <a href="https://github.com/AlexR2D2/metabase_duckdb_driver/pull/19">PR</a> for this will land very soon giving a good quick way to link to the latest and greatest in duckdb from metabase</p>
|
||||||
|
<h3>But How do we get Data?</h3>
|
||||||
|
<p>Brilliant, using the recomended DockerFile we can load up a metabase container with the duckdb driver pre built</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">46.2</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">AlexR2D2</span><span class="o">/</span><span class="n">metabase_duckdb_driver</span><span class="o">/</span><span class="n">releases</span><span class="o">/</span><span class="n">download</span><span class="o">/</span><span class="mf">0.1</span><span class="o">.</span><span class="mi">6</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;java&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;-jar&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/metabase.jar&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>Great Now the big question. How do we get the data into the damn thing. Interestingly initially when I was designing this I had the thought of leveraging the in memory capabilities of duckdb and pulling in from the parquet on s3 directly as needed, after all the cluster is on AWS so the s3 API requests should be unbelievably fast anyway so why bother with a persistent database? </p>
|
||||||
|
<p>Now that we have the default credentials chain it is trivial to call parquet from s3</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">SELECT</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="k">FROM</span><span class="w"> </span><span class="n">read_parquet</span><span class="p">(</span><span class="s1">&#39;s3://&lt;bucket&gt;/&lt;file&gt;&#39;</span><span class="p">);</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>However, if you're reading direct off parquet all of a sudden you need to consider the partioning and I also found out that, if the parquet is being actively written to at the time of quering, duckdb has a hissyfit about metadata not matching the query. Needless to say duckdb and streaming parquet are not happy bed fellows (<em>and frankly were not desined to be so this is ok</em>). And the idea of trying to explain all this to the run of the mill reporting analyst whom it is my hope is a business sort of person not tech honestly gave me hives.. so I had to make it easier</p>
|
||||||
|
<p>The compromise occured to me... the curated layer is only built daily for reporting, and using that, I could create a duckdb file on disk that could be loaded into the metabase container itself.</p>
|
||||||
|
<p>With some very simple python as an operation in our orchestrator I had a job that would read direct from our curated parquet and create a duckdb file with it.. without giving away to much the job primarily consisted of this </p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">def</span> <span class="nf">duckdb_builder</span><span class="p">(</span><span class="n">table</span><span class="p">):</span>
|
||||||
|
<span class="n">conn</span> <span class="o">=</span> <span class="n">duckdb</span><span class="o">.</span><span class="n">connect</span><span class="p">(</span><span class="s2">&quot;curated_duckdb.duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;CALL load_aws_credentials(&#39;</span><span class="si">{</span><span class="n">aws_profile</span><span class="si">}</span><span class="s2">&#39;)&quot;</span><span class="p">)</span>
|
||||||
|
<span class="c1">#This removes a lot of weirdass ANSI in logs you DO NOT WANT</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">execute</span><span class="p">(</span><span class="s2">&quot;PRAGMA enable_progress_bar=false&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;Create </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> in duckdb&quot;</span><span class="p">)</span>
|
||||||
|
<span class="n">sql</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">&quot;CREATE OR REPLACE TABLE </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> AS SELECT * FROM read_parquet(&#39;s3://</span><span class="si">{</span><span class="n">curated_bucket</span><span class="si">}</span><span class="s2">/</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2">/*&#39;)&quot;</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="n">sql</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">&quot;</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> Created&quot;</span><span class="p">)</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And then an upload to an s3 bucket</p>
|
||||||
|
<p>This of course necessated a cron job baked in to the metabase container itself to actually pull the duckdb in every morning. After some carefuly analysis of time (because I'm do lazy to implement message queues) I set up a s3 cp job that could be cronned direct from the container itself. This gives us a self updating metabase container pulling with a duckdb backend for client facing reporting right in the interface. AND because of the fact the duckdb is baked right into the container... there are NO associated s3 or dpu costs (merely the cost of running a relatively large container)</p>
|
||||||
|
<p>The final Dockerfile looks like this</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">47.6</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">mkdir</span><span class="w"> </span><span class="o">-</span><span class="n">p</span><span class="w"> </span><span class="o">/</span><span class="n">duckdb_data</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">entrypoint</span><span class="o">.</span><span class="n">sh</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">helper_scripts</span><span class="o">/</span><span class="n">download_duckdb</span><span class="o">.</span><span class="n">py</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">update</span><span class="w"> </span><span class="o">-</span><span class="n">y</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">upgrade</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">python3</span><span class="w"> </span><span class="n">python3</span><span class="o">-</span><span class="n">pip</span><span class="w"> </span><span class="n">cron</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">pip3</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">boto3</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span><span class="n">l</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">cat</span><span class="p">;</span><span class="w"> </span><span class="n">echo</span><span class="w"> </span><span class="s2">&quot;0 */6 * * * python3 /home/helper_scripts/download_duckdb.py&quot;</span><span class="p">;</span><span class="w"> </span><span class="p">}</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">&quot;bash&quot;</span><span class="p">,</span><span class="w"> </span><span class="s2">&quot;/home/entrypoint.sh&quot;</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And there we have it... an in memory containerised reporting solution with blazing fast capability to aggregate and build reports based on curated data direct from the business.. fully automated and deployable via CI/CD, that provides data updates daily.</p>
|
||||||
|
<p>Now the embedded part.. which isn't built yet but I'll make sure to update you once we have/if we do because the architecture is very exciting for an embbdedded reporting workflow that is deployable via CI/CD processes to applications. As a little taster I'll point you to the <a href="https://www.metabase.com/learn/administration/git-based-workflow">metabase documentation</a>, the unfortunate thing about it is Metabase <em>have</em> hidden this behind the enterprise license.. but I can absolutely see why. If we get to implementing this I'll be sure to update you here on the learnings.</p>
|
||||||
|
<p>Until then....</p></content><category term="Business Intelligence"></category><category term="data engineering"></category><category term="Metabase"></category><category term="DuckDB"></category><category term="embedded"></category></entry></feed>
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Data Analytics</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/data-analytics.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-07-13T20:00:00+10:00</updated><entry><title>Notebook or BI, What is the most appropiate communication medium</title><link href="http://localhost:8000/notebook-or-bi.html" rel="alternate"></link><published>2023-07-13T20:00:00+10:00</published><updated>2023-07-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-07-13:/notebook-or-bi.html</id><summary type="html"><p>When is a notebook enough or when do we need a dashboard</p></summary><content type="html"><p>I want to preface this post by saying I think "Dashboards" or "BI" as terms are wayyyyyyyyyyyyyyyyy over saturated in the market. There seems to be a belief that any question answerable in data deserves the work associated with a dashboard when in fact a simple one off report, or notebook, would be more than enough.</p></content><category term="Data Analytics"></category><category term="data engineering"></category><category term="Data Analytics"></category></entry></feed>
|
||||||
@@ -1,54 +1,5 @@
|
|||||||
<?xml version="1.0" encoding="utf-8"?>
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Data Engineering</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/data-engineering.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-06-15T20:00:00+10:00</updated><entry><title>CI/CD in Data Engineering</title><link href="http://localhost:8000/CI/CD%20in%20Data%20and%20Data%20Infrastructure.html" rel="alternate"></link><published>2023-06-15T20:00:00+10:00</published><updated>2023-06-15T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-06-15:/CI/CD in Data and Data Infrastructure.html</id><summary type="html"><p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p></summary><content type="html"><p>Data Engineering has traditionally been considered the bastard step child of work that would have once been considered Administrative in the Tech world. Predominately we write SQL and then deploy that SQL onto one or more Databases. In fact a lot of the traditional methodologies around data almost assume this is the core of how an organistation is managing the majority of it's data. In the last couple of years though there has been a very steady move towards having the Data Engineering workload of SQL move towards Software Engineering techniques. With the popularity of tools like <a href="https://www.dbtlabs.com">DBT</a> and the latest newcommer on the block, <a href="https://www.sqlmesh.com">SQL-MESH</a> The oppportunity has started to arise where we can align our Data Engineering workloads with different environments and move much more efficiently towards a Continous Integration and Deployment methodology in our workflows. </p>
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Data Engineering</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/data-engineering.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2023-05-23T20:00:00+10:00</updated><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
||||||
<p>For the Data Engineering space the move to the cloud has been a breath of fresh air (Not so in some other IT disciplines). I am relatively young, so I don't 100% remember but my experience has taught me that there were 3 options here not so long ago</p>
|
|
||||||
<p><em>Expensive:</em></p>
|
|
||||||
<ul>
|
|
||||||
<li>SAS</li>
|
|
||||||
<li>SSIS/SSRS</li>
|
|
||||||
<li>COGNOS/TM1</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>Rickety:</em></p>
|
|
||||||
<ul>
|
|
||||||
<li>Just write stored procedures!</li>
|
|
||||||
<li>Startup script on my laptop XD</li>
|
|
||||||
<li>"Don't touch that machine over there, No one knows what it does but if it's turned off our financial reports don't work" (This is a third hand story I heard, seriously!)</li>
|
|
||||||
<li>"I need to an upgrade to my laptop, Excel needs more than 8GB of RAM"</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>Hard:</em></p>
|
|
||||||
<ul>
|
|
||||||
<li>Hadoop</li>
|
|
||||||
<li>Spark (hadoop but whatever)</li>
|
|
||||||
<li>Python</li>
|
|
||||||
<li>R</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>(The reason I've listed them as hard is because self hosting Hadoop/Spark and managing a truckload of python or R scripts, whilst it could have been "cheap" required a team of devs who really really knew what they were doing... so not really cheap and also <strong>really hard</strong>)</em></p>
|
|
||||||
<p>Then there was getting git behind all the sql scripts and modelling, let alone CI/CD <strong>IF</strong> it existed, it was custom, and bespoke and likely had a single point of failure in the person who knew how <code>git merge</code> worked. At least... thats I was told I'm not <em>that</em> old ;p.</p>
|
|
||||||
<p>These days we are pretty blessed, with democrotisation of clusters and data Infrastructure in the cloud we no longer need a team of sysadmins who know how to tune a cluster to the Nth degree to get the best our of our data workloads (well... we do, but we pay the cloud guys for that!). However, we still need to know about the idiosyncracities of this infrastructure, when it is appropiate to use and how we want to control and maintain the workloads. </p>
|
|
||||||
<p>In general when I am designing a system I normally like to break it into 3.</p>
|
|
||||||
<ul>
|
|
||||||
<li>Storage</li>
|
|
||||||
<li>Compute</li>
|
|
||||||
<li>Code</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>In General</em> Storage and Compute will be infrastructer related, "Code" is sort of a catch all for my modelling, normally sql, python/r or spark scripts that are used to provide system or business logic, anything really thats going to get data to the end user/analyst/data scientists/annoying person </p>
|
|
||||||
<p>Traditionally the compute layer only really had 2 considerations</p>
|
|
||||||
<ul>
|
|
||||||
<li>SQL or Logic engine (normally a flavour of spark(glue) and then something like reshift/athena/trino/bigquery)</li>
|
|
||||||
<li>Orchestration Layer (Airflow, Dagster)</li>
|
|
||||||
</ul>
|
|
||||||
<p>But with the advent of sql engine agnostic Modelling we potentially now need to also consider</p>
|
|
||||||
<ul>
|
|
||||||
<li>Model Compilation</li>
|
|
||||||
</ul>
|
|
||||||
<p>Now on the surface it seems counterintuitive to seperate the models from the logic layer but lets consider the following scenario</p>
|
|
||||||
<blockquote>
|
|
||||||
<p>Redshift is costing to much and is getting slow, we want to try bigquery
|
|
||||||
How much investment will it be to change over</p>
|
|
||||||
</blockquote>
|
|
||||||
<p>Now, If the entirety of your modelling is stored and deployed to big query direct this would not only involve the investment of spinning up the big query account and either connecting or migrating your data over. You would also need to consider how in the bloody hell you convert all your existing models and workflows over.</p>
|
|
||||||
<p>With something like DBT or sqlmesh you change your compilation target and it's done for you. It also means the Data Engineer now doesn't need to necessarily understand the esoteric nature of the target, at least for simple models (which, lets be real, most are).</p>
|
|
||||||
<p>BUT, now we have <em>a lot</em> of software and infrastructure a simplified common datastack will look something like the below (Assuming ELT, ETL is a bit different but more or less needs the same components)</p>
|
|
||||||
<p><img src="http://localhost:8000/images/DataStackSimplified.png" width="600" height="295" /></p></content><category term="Data Engineering"></category><category term="data engineering"></category><category term="DBT"></category><category term="Terraform"></category><category term="IAC"></category></entry><entry><title>Implmenting Appflow in a Production Datalake</title><link href="http://localhost:8000/appflow-production.html" rel="alternate"></link><published>2023-05-23T20:00:00+10:00</published><updated>2023-05-17T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2023-05-23:/appflow-production.html</id><summary type="html"><p>How Appflow simplified a major extract layer and when I choose Managed Services</p></summary><content type="html"><p>I recently attended a meetup where there was a talk by an AWS spokesperson. Now don't get me wrong, I normally take these things with a grain of salt. At this talk there was this tiny tiny little segment about a product that AWS had released called <a href="https://aws.amazon.com/appflow/">Amazon Appflow</a>. This product <em>claimed</em> to be able to automate and make easy the link between different API endpoints, REST or otherwise and send that data to another point, whether that is Redshift, Aurora, a general relational db in RDS or otherwise or s3.</p>
|
|
||||||
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
<p>This was particularly interesting to me because I had recently finished creating and s3 datalake in AWS for the company I work for. Today, I finally put my first Appflow integration to the Datalake into production and I have to say there are some rough edges to the deployment but it has been more or less as described on the box. </p>
|
||||||
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
<p>Over the course of the next few paragraphs I'd like to explain the thinking I had as I investigated the product and then ultimately why I chose a managed service for this over implementing something myself in python using Dagster which I have also spun up within our cluster on AWS.</p>
|
||||||
<h3>Datalake Extraction Layer</h3>
|
<h3>Datalake Extraction Layer</h3>
|
||||||
|
|||||||
@@ -0,0 +1,143 @@
|
|||||||
|
<?xml version="1.0" encoding="utf-8"?>
|
||||||
|
<feed xmlns="http://www.w3.org/2005/Atom"><title>Andrew Ridgway's Blog - Resume</title><link href="http://localhost:8000/" rel="alternate"></link><link href="http://localhost:8000/feeds/resume.atom.xml" rel="self"></link><id>http://localhost:8000/</id><updated>2024-03-13T20:00:00+10:00</updated><entry><title>A Cover Letter</title><link href="http://localhost:8000/cover-letter.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/cover-letter.html</id><summary type="html"><p>A Summary of what I've done and Where I'd like to go for prospective Employers</p></summary><content type="html"><p>To whom it may concern</p>
|
||||||
|
<p>My name is Andrew Ridgway and I am a Data and Technology professional looking to embark on the next step in my career.</p>
|
||||||
|
<p>I have over 10 years’ experience in System and Data Architecture, Data Modelling and Orchestration, Business and Technical Analysis and System and Development Process Design. Most of this has been in developing Cloud architectures and workloads on AWS and GCP Including ML workloads using Sagemaker. </p>
|
||||||
|
<p>In my current role I have Proposed, Designed and built the data platform currently used by business. This includes internal and external data products as well as the infrastructure and modelling to support these. This role has seen me liaise with stakeholders of all levels of the business from Analysts in the Customer Experience team right up to C suite executives and preparing material for board members. I understand the complexity of communicating complex system design to different level stakeholders and the complexities of involved in communicating to both technical and less technical employees particularly in relation to data and ML technologies. </p>
|
||||||
|
<p>I have also worked as a technical consultant to many businesses and have assisted with the design and implementation of systems for a wide range of industries including financial services, mining and retail. I understand the complexities created by regulation in these environments and understand that this can sometimes necessitate the use of technologies and designs, including legacy systems and designs, I wouldn’t normally use. I also have a passion of designing systems that enable these organisations to realise the benefits of CI/CD on workloads they would not traditionally use this capability. In particular I took a very traditional legacy Data Warehousing team and implemented a solution that meant version control was no longer controlled by a daily copy and paste of folders with dates on major updates. My solution involved establishing guidelines of use of git version control so that this could happen automatically as people committed new code to the core code base. As I have moved into cloud architecture I have made sure to use best practice and ensure everything I build isn’t considered production ready until it is in IAC and deployed through a CI/CD pipeline.</p>
|
||||||
|
<p>In a personal capacity I am an avid tech and ML enthusiast. I have designed my own cluster including monitoring and deployment that runs several services that my family uses including chat and DNS and am in the process of designing a “set and forget” system that will allows me to have multi user tenancies on hardware I operate that should enable us to have the niceties of cloud services like email, storage and scheduling with the safety of knowing where that data is stored and exactly how it is used. I also like to design small IoT devices out of Arduino boards allowing me to monitor and control different facets of our house like temperature and light. </p>
|
||||||
|
<p>Currently I am working on a project to merge my skill in SQL Modelling and Orchestration with GPT API’s to try and lessen that burden. You can see some of this work in its very early stages here:
|
||||||
|
(gpt-sql-generator)[https://github.com/armistace/gpt-sql-generator]
|
||||||
|
(dbt_sources_generator)[https://github.com/armistace/datahub_dbt_sources_generator]</p>
|
||||||
|
<p>I look forward to hearing from you soon.</p>
|
||||||
|
<p>Sincerely,</p>
|
||||||
|
<hr>
|
||||||
|
<p>Andrew Ridgway</p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry><entry><title>A Resume</title><link href="http://localhost:8000/resume.html" rel="alternate"></link><published>2024-02-23T20:00:00+10:00</published><updated>2024-03-13T20:00:00+10:00</updated><author><name>Andrew Ridgway</name></author><id>tag:localhost,2024-02-23:/resume.html</id><summary type="html"><p>A Summary of My work Experience</p></summary><content type="html"><h1>OVERVIEW</h1>
|
||||||
|
<p>I am a Senior Data Engineer looking to transition my skills to Data and Solution
|
||||||
|
Architecting as well as project management. I have spent the better part of the
|
||||||
|
last decade refining my abilities in taking business requirements and turning
|
||||||
|
those into actionable data engineering, analytics, and software projects with
|
||||||
|
trackable metrics. I believe in agnosticism when it comes to coding languages
|
||||||
|
and have experimented in my own time with many different languages. In my
|
||||||
|
career I have used Python, .NET, PowerShell, TSQL, VB and SAS (multiple
|
||||||
|
products) in an Enterprise capacity. I also have experience using Google Cloud
|
||||||
|
Platform and AWS tools for ETL and data platform development as well as git
|
||||||
|
for version control and deployment using various IAC tools. I have also
|
||||||
|
conducted data analysis and modelling on business metrics to find relationships
|
||||||
|
between both staff and customer behavior and produced actionable
|
||||||
|
recommendations based on the conclusions. In a private context I have also
|
||||||
|
experimented with C, C# and Kotlin I am looking to further my career by taking
|
||||||
|
my passion for data engineering and analysis as well as web and software
|
||||||
|
development and applying it in a strategic context.</p>
|
||||||
|
<h1>SKILLS &amp; ABILITIES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Python (scripting, compiling, notebooks – Sagemaker, Jupyter)</li>
|
||||||
|
<li>git</li>
|
||||||
|
<li>SAS (Base, EG, VA)</li>
|
||||||
|
<li>Various Google Cloud Tools (Data Fusion, Compute Engine, Cloud Functions)</li>
|
||||||
|
<li>Various Amazon Tools (EC2, RDS, Kinesis, Glue, Redshift, Lambda, ECS, ECR, EKS)</li>
|
||||||
|
<li>Streaming Technologies (Kafka, Hive, Spark Streaming)</li>
|
||||||
|
<li>Various DB platforms both on Prem and Serverless (MariaDB/MySql,</li>
|
||||||
|
<li>Postgres/Redshift, SQL Server, RDS/Aurora variants)</li>
|
||||||
|
<li>Various Microsoft Products (PowerBI, TSQL, Excel, VBA)</li>
|
||||||
|
<li>Linux Server Administration (cron, bash, systemD)</li>
|
||||||
|
<li>ETL/ELT Development</li>
|
||||||
|
<li>Basic Data Modelling (Kimball, SCD Type 2)</li>
|
||||||
|
<li>IAC (Cloud Formation, Terraform)</li>
|
||||||
|
<li>Datahub Deployment</li>
|
||||||
|
<li>Dagster Orchestration Deployments</li>
|
||||||
|
<li>DBT Modelling and Design Deployments</li>
|
||||||
|
<li>Containerised and Cloud Driven Data Architecture</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EXPERIENCE</h1>
|
||||||
|
<h2>Cloud Data Architect</h2>
|
||||||
|
<h3><em>Redeye Apps</em></h3>
|
||||||
|
<h4><em>May 2022 - Present</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Greenfields Research, Design and Deployment of S3 datalake (Parquet)</li>
|
||||||
|
<li>AWS DMS, S3, Athena, Glue</li>
|
||||||
|
<li>Research Design and Deployment of Catalog (Datahub)</li>
|
||||||
|
<li>Design of Data Governance Process (Datahub driven)</li>
|
||||||
|
<li>Research Design and Deployment of Orchestration and Modelling for Transforms (Dagster/DBT into Mesos)</li>
|
||||||
|
<li>CI/CD design and deployment of modelling and orchestration using Gitlab</li>
|
||||||
|
<li>Research, Design and Deployment of ML Ops Dev pipelines anddeployment strategy</li>
|
||||||
|
<li>Design of ETL/Pipelines (DBT)</li>
|
||||||
|
<li>Design of Customer Facing Data Products and deployment methodologies (Fully automated via Kakfa/Dagster/DBT)</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Data Engineer,</h2>
|
||||||
|
<h3><em>TechConnect IT Solutions</em></h3>
|
||||||
|
<h4><em>August 2021 – May 2022</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Design of Cloud Data Batch ETL solutions using Python (Glue)</li>
|
||||||
|
<li>Design of Cloud Data Streaming ETL solution using Python (Kinesis)</li>
|
||||||
|
<li>Solve complex client business problems using software to join and transform data from DB’s, Web API’s, Application API’s and System logs</li>
|
||||||
|
<li>Build CI/CD pipelines to ensure smooth deployments (Bitbucket, gitlab)</li>
|
||||||
|
<li>Apply Prebuilt ML models to software solutions (Sagemaker)</li>
|
||||||
|
<li>Assist with the architecting of Containerisation solutions (Docker, ECS, ECR)</li>
|
||||||
|
<li>API testing and development (gRPC, Rest)</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Enterprise Data Warehouse Developer</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>August 2019 - August 2021</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>ETL development of CRM, WFP, Outbound Dialer, Inbound switch in Google Cloud, SAS, TSQL</li>
|
||||||
|
<li>Bringing new data to the business to analyse for new insights</li>
|
||||||
|
<li>Redeveloped Version Control and brought git to the data team</li>
|
||||||
|
<li>Introduced python for API enablement in the Enterprise Data Warehouse</li>
|
||||||
|
<li>Partnering with the business to focus data project on actual need and translating into technical requirements</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Business Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2018 - August 2019</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Automate Service Performance Reporting using PowerShell/VBA/SAS</li>
|
||||||
|
<li>Learn and leverage SAS EG and VA to streamline Microsoft Excel Reporting</li>
|
||||||
|
<li>Identify and develop data pipelines to source data from multiple sources easily and collate into a single source to identify relationships and trends</li>
|
||||||
|
<li>Technologies used include VBA, PowerShell, SQL, Web API’s, SAS</li>
|
||||||
|
<li>Where SAS is inappropriate use VBA to automate processes in Microsoft Access and Excel</li>
|
||||||
|
<li>Gather Requirements to build meaningful reporting solutions</li>
|
||||||
|
<li>Provide meaningful analysis on business performance and provide relevant presentations and reports to senior stakeholders.</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Forecasting and Capacity Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2017 – January 2018</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop the outbound forecasting model for the Auto and General sales call center by analysing the relationship between customer decisions and workload drivers</li>
|
||||||
|
<li>This includes the complete data pipeline for the model from identifying and sourcing data, building the reporting and analysing the data and associated drivers.</li>
|
||||||
|
<li>Forecast inbound workload requirements for the Auto and General sales call center using time series analysis</li>
|
||||||
|
<li>Learn and leverage the Aspect Workforce Management System to ensure efficiency of forecast generation</li>
|
||||||
|
<li>Learn and leverage the capabilities of SAS Enterprise Guide to improve accuracy</li>
|
||||||
|
<li>Liaise with people across the business to ensure meaningful, accurate analysis is provided to senior stakeholders</li>
|
||||||
|
<li>Analyse monthly, weekly and intraday requirements and ensure forecast is accurately predicting workload for breaks, meetings and Leave</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Senior HR Performance Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>June 2016 - January 2017</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Harmonise various systems to develop a unified workforce reporting and analysis framework with appropriate metrics</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Workforce Business Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>July 2015 – June 2016</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop and refine current workforce analysis techniques and databases</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Act as liaison between shared service providers and executives and facilitate communication during the implementation of a payroll leave audit</li>
|
||||||
|
<li>Gather reporting requirements from various business areas and produce ad-hoc and regular reports as required</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EDUCATION</h1>
|
||||||
|
<ul>
|
||||||
|
<li>2011 Bachelor of Business Management, University of Queensland</li>
|
||||||
|
<li>2008 Bachelor of Arts, University of Queensland</li>
|
||||||
|
</ul>
|
||||||
|
<h1>REFERENCES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Anthony Stiller Lead Developer, Data warehousing, Queensland Health</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0428 038 031</em></p>
|
||||||
|
<ul>
|
||||||
|
<li>Jaime Brian Head of Cloud Ninjas, TechConnect</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0422 012 17</em></p></content><category term="Resume"></category><category term="Cover Letter"></category><category term="Resume"></category></entry></feed>
|
||||||
Binary file not shown.
|
Before Width: | Height: | Size: 323 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 146 KiB |
+30
-4
@@ -85,15 +85,41 @@
|
|||||||
<div class="row">
|
<div class="row">
|
||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<div class="post-preview">
|
<div class="post-preview">
|
||||||
<a href="http://localhost:8000/CI/CD in Data and Data Infrastructure.html" rel="bookmark" title="Permalink to CI/CD in Data Engineering">
|
<a href="http://localhost:8000/cover-letter.html" rel="bookmark" title="Permalink to A Cover Letter">
|
||||||
<h2 class="post-title">
|
<h2 class="post-title">
|
||||||
CI/CD in Data Engineering
|
A Cover Letter
|
||||||
</h2>
|
</h2>
|
||||||
</a>
|
</a>
|
||||||
<p>When to use IaC CI/CD techniques or Software CI/CD techniques in Data Architecture</p>
|
<p>A Summary of what I've done and Where I'd like to go for prospective Employers</p>
|
||||||
<p class="post-meta">Posted by
|
<p class="post-meta">Posted by
|
||||||
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
on Thu 15 June 2023
|
on Fri 23 February 2024
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/resume.html" rel="bookmark" title="Permalink to A Resume">
|
||||||
|
<h2 class="post-title">
|
||||||
|
A Resume
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>A Summary of My work Experience</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Fri 23 February 2024
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
<hr>
|
||||||
|
<div class="post-preview">
|
||||||
|
<a href="http://localhost:8000/metabase-duckdb.html" rel="bookmark" title="Permalink to Metabase and DuckDB">
|
||||||
|
<h2 class="post-title">
|
||||||
|
Metabase and DuckDB
|
||||||
|
</h2>
|
||||||
|
</a>
|
||||||
|
<p>Using Metabase and DuckDB to create an embedded Reporting Container bringing the data as close to the report as possible</p>
|
||||||
|
<p class="post-meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Wed 15 November 2023
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
<hr>
|
<hr>
|
||||||
|
|||||||
@@ -0,0 +1,246 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<meta name="description" content="">
|
||||||
|
<meta name="author" content="">
|
||||||
|
|
||||||
|
<title>Andrew Ridgway's Blog</title>
|
||||||
|
|
||||||
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
|
<link href="http://localhost:8000/feeds/business-intelligence.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
|
<!-- Bootstrap Core CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/clean-blog.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Code highlight color scheme -->
|
||||||
|
<link href="http://localhost:8000/theme/css/code_blocks/tomorrow.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom Fonts -->
|
||||||
|
<link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'>
|
||||||
|
|
||||||
|
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
|
||||||
|
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
|
||||||
|
<!--[if lt IE 9]>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
|
||||||
|
<![endif]-->
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
<meta name="tags" contents="data engineering" />
|
||||||
|
<meta name="tags" contents="Metabase" />
|
||||||
|
<meta name="tags" contents="DuckDB" />
|
||||||
|
<meta name="tags" contents="embedded" />
|
||||||
|
|
||||||
|
|
||||||
|
<meta property="og:locale" content="en">
|
||||||
|
<meta property="og:site_name" content="Andrew Ridgway's Blog">
|
||||||
|
|
||||||
|
<meta property="og:type" content="article">
|
||||||
|
<meta property="article:author" content="">
|
||||||
|
<meta property="og:url" content="http://localhost:8000/metabase-duckdb.html">
|
||||||
|
<meta property="og:title" content="Metabase and DuckDB">
|
||||||
|
<meta property="og:description" content="">
|
||||||
|
<meta property="og:image" content="http://localhost:8000/">
|
||||||
|
<meta property="article:published_time" content="2023-11-15 20:00:00+10:00">
|
||||||
|
</head>
|
||||||
|
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<!-- Navigation -->
|
||||||
|
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
|
||||||
|
<div class="container-fluid">
|
||||||
|
<!-- Brand and toggle get grouped for better mobile display -->
|
||||||
|
<div class="navbar-header page-scroll">
|
||||||
|
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
|
||||||
|
<span class="sr-only">Toggle navigation</span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
</button>
|
||||||
|
<a class="navbar-brand" href="http://localhost:8000/">Andrew Ridgway's Blog</a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Collect the nav links, forms, and other content for toggling -->
|
||||||
|
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
|
||||||
|
<ul class="nav navbar-nav navbar-right">
|
||||||
|
|
||||||
|
</ul>
|
||||||
|
</div>
|
||||||
|
<!-- /.navbar-collapse -->
|
||||||
|
</div>
|
||||||
|
<!-- /.container -->
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<!-- Page Header -->
|
||||||
|
<header class="intro-header" style="background-image: url('http://localhost:8000/theme/images/post-bg.jpg')">
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-heading">
|
||||||
|
<h1>Metabase and DuckDB</h1>
|
||||||
|
<span class="meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Wed 15 November 2023
|
||||||
|
</span>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- Main Content -->
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<!-- Post Content -->
|
||||||
|
<article>
|
||||||
|
<p>Ahhhh <a href="https://duckdb.org/">DuckDB</a> if you're even partly floating around in the data space you've probably been hearing ALOT about it and it's <em>"Datawarehouse on your laptop"</em> mantra. However, the OTHER application that sometimes gets missed is <em>"SQLite for OLAP workloads"</em> and it was this concept that once I grasped it gave me a very interesting idea.... What if we could take the very pretty Aggregate Layer of our Data(warehouse/LakeHouse/Lake) and put that data right next to presentation layer of the lake, reducing network latency and... hopefully... have presentation reports running over very large workloads in the blink of an eye. It might even be fast enough that it could be deployed and embedded </p>
|
||||||
|
<p>However, for this to work we need some form of conatinerised reporting application.... lucky for us there is <a href="https://www.metabase.com/">Metabase</a> which is a fantastic little reporting application that has an open core. So this got me thinking... Can I put these two applications together and create a Reporting Layer with report embedding capabilities that is deployable in the cluster and has a admin UI accesible over a web page all whilst keeping the data locked to our network?</p>
|
||||||
|
<h3>The Beginnings of an Idea</h3>
|
||||||
|
<p>Ok so... Big first question. Can Duckdb and Metabase talk? Well... not quite. But first lets take a quick look at the architecture we'll be employing here </p>
|
||||||
|
<p><img alt="Duckdb Architecture" height="auto" width="100%" src="http://localhost:8000/images/metabase_duckdb.png"></p>
|
||||||
|
<p>But you'll notice this pretty glossed over line, "Connector", that right there is the clincher. So what is this "Connector"?. </p>
|
||||||
|
<p>To Deep dive into this would take a whole blog so to give you something to quickly wrap your head around its the glue that will make metabase be able to query your data source. The reality is its a jdbc driver compiled against metabase. </p>
|
||||||
|
<p>Thankfully Metabase point you to a <a href="https://github.com/AlexR2D2/metabase_duckdb_driver">community driver</a> for linking to duckdb ( hopefully it will be brought into metabase proper sooner rather than later ) </p>
|
||||||
|
<p>Now the release of this driver is still compiled against 0.8 of duckdb and 0.9 is the latest stable but hopefully the <a href="https://github.com/AlexR2D2/metabase_duckdb_driver/pull/19">PR</a> for this will land very soon giving a good quick way to link to the latest and greatest in duckdb from metabase</p>
|
||||||
|
<h3>But How do we get Data?</h3>
|
||||||
|
<p>Brilliant, using the recomended DockerFile we can load up a metabase container with the duckdb driver pre built</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">46.2</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">github</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">AlexR2D2</span><span class="o">/</span><span class="n">metabase_duckdb_driver</span><span class="o">/</span><span class="n">releases</span><span class="o">/</span><span class="n">download</span><span class="o">/</span><span class="mf">0.1</span><span class="o">.</span><span class="mi">6</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">"java"</span><span class="p">,</span><span class="w"> </span><span class="s2">"-jar"</span><span class="p">,</span><span class="w"> </span><span class="s2">"/home/metabase.jar"</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>Great Now the big question. How do we get the data into the damn thing. Interestingly initially when I was designing this I had the thought of leveraging the in memory capabilities of duckdb and pulling in from the parquet on s3 directly as needed, after all the cluster is on AWS so the s3 API requests should be unbelievably fast anyway so why bother with a persistent database? </p>
|
||||||
|
<p>Now that we have the default credentials chain it is trivial to call parquet from s3</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">SELECT</span><span class="w"> </span><span class="o">*</span><span class="w"> </span><span class="k">FROM</span><span class="w"> </span><span class="n">read_parquet</span><span class="p">(</span><span class="s1">'s3://<bucket>/<file>'</span><span class="p">);</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>However, if you're reading direct off parquet all of a sudden you need to consider the partioning and I also found out that, if the parquet is being actively written to at the time of quering, duckdb has a hissyfit about metadata not matching the query. Needless to say duckdb and streaming parquet are not happy bed fellows (<em>and frankly were not desined to be so this is ok</em>). And the idea of trying to explain all this to the run of the mill reporting analyst whom it is my hope is a business sort of person not tech honestly gave me hives.. so I had to make it easier</p>
|
||||||
|
<p>The compromise occured to me... the curated layer is only built daily for reporting, and using that, I could create a duckdb file on disk that could be loaded into the metabase container itself.</p>
|
||||||
|
<p>With some very simple python as an operation in our orchestrator I had a job that would read direct from our curated parquet and create a duckdb file with it.. without giving away to much the job primarily consisted of this </p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="k">def</span> <span class="nf">duckdb_builder</span><span class="p">(</span><span class="n">table</span><span class="p">):</span>
|
||||||
|
<span class="n">conn</span> <span class="o">=</span> <span class="n">duckdb</span><span class="o">.</span><span class="n">connect</span><span class="p">(</span><span class="s2">"curated_duckdb.duckdb"</span><span class="p">)</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="sa">f</span><span class="s2">"CALL load_aws_credentials('</span><span class="si">{</span><span class="n">aws_profile</span><span class="si">}</span><span class="s2">')"</span><span class="p">)</span>
|
||||||
|
<span class="c1">#This removes a lot of weirdass ANSI in logs you DO NOT WANT</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">execute</span><span class="p">(</span><span class="s2">"PRAGMA enable_progress_bar=false"</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">"Create </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> in duckdb"</span><span class="p">)</span>
|
||||||
|
<span class="n">sql</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">"CREATE OR REPLACE TABLE </span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> AS SELECT * FROM read_parquet('s3://</span><span class="si">{</span><span class="n">curated_bucket</span><span class="si">}</span><span class="s2">/</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2">/*')"</span>
|
||||||
|
<span class="n">conn</span><span class="o">.</span><span class="n">sql</span><span class="p">(</span><span class="n">sql</span><span class="p">)</span>
|
||||||
|
<span class="n">log</span><span class="o">.</span><span class="n">info</span><span class="p">(</span><span class="sa">f</span><span class="s2">"</span><span class="si">{</span><span class="n">table</span><span class="si">}</span><span class="s2"> Created"</span><span class="p">)</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And then an upload to an s3 bucket</p>
|
||||||
|
<p>This of course necessated a cron job baked in to the metabase container itself to actually pull the duckdb in every morning. After some carefuly analysis of time (because I'm do lazy to implement message queues) I set up a s3 cp job that could be cronned direct from the container itself. This gives us a self updating metabase container pulling with a duckdb backend for client facing reporting right in the interface. AND because of the fact the duckdb is baked right into the container... there are NO associated s3 or dpu costs (merely the cost of running a relatively large container)</p>
|
||||||
|
<p>The final Dockerfile looks like this</p>
|
||||||
|
<div class="highlight"><pre><span></span><code><span class="n">FROM</span><span class="w"> </span><span class="n">openjdk</span><span class="p">:</span><span class="mi">19</span><span class="o">-</span><span class="n">buster</span>
|
||||||
|
|
||||||
|
<span class="n">ENV</span><span class="w"> </span><span class="n">MB_PLUGINS_DIR</span><span class="o">=/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">https</span><span class="p">:</span><span class="o">//</span><span class="n">downloads</span><span class="o">.</span><span class="n">metabase</span><span class="o">.</span><span class="n">com</span><span class="o">/</span><span class="n">v0</span><span class="o">.</span><span class="mf">47.6</span><span class="o">/</span><span class="n">metabase</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
<span class="n">ADD</span><span class="w"> </span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">chmod</span><span class="w"> </span><span class="mi">744</span><span class="w"> </span><span class="o">/</span><span class="n">home</span><span class="o">/</span><span class="n">plugins</span><span class="o">/</span><span class="n">duckdb</span><span class="o">.</span><span class="n">metabase</span><span class="o">-</span><span class="n">driver</span><span class="o">.</span><span class="n">jar</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">mkdir</span><span class="w"> </span><span class="o">-</span><span class="n">p</span><span class="w"> </span><span class="o">/</span><span class="n">duckdb_data</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">entrypoint</span><span class="o">.</span><span class="n">sh</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">COPY</span><span class="w"> </span><span class="n">helper_scripts</span><span class="o">/</span><span class="n">download_duckdb</span><span class="o">.</span><span class="n">py</span><span class="w"> </span><span class="o">/</span><span class="n">home</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">update</span><span class="w"> </span><span class="o">-</span><span class="n">y</span><span class="w"> </span><span class="o">&&</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">upgrade</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">apt</span><span class="o">-</span><span class="n">get</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">python3</span><span class="w"> </span><span class="n">python3</span><span class="o">-</span><span class="n">pip</span><span class="w"> </span><span class="n">cron</span><span class="w"> </span><span class="o">-</span><span class="n">y</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">pip3</span><span class="w"> </span><span class="n">install</span><span class="w"> </span><span class="n">boto3</span>
|
||||||
|
|
||||||
|
<span class="n">RUN</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span><span class="n">l</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">cat</span><span class="p">;</span><span class="w"> </span><span class="n">echo</span><span class="w"> </span><span class="s2">"0 */6 * * * python3 /home/helper_scripts/download_duckdb.py"</span><span class="p">;</span><span class="w"> </span><span class="p">}</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="n">crontab</span><span class="w"> </span><span class="o">-</span>
|
||||||
|
|
||||||
|
<span class="n">CMD</span><span class="w"> </span><span class="p">[</span><span class="s2">"bash"</span><span class="p">,</span><span class="w"> </span><span class="s2">"/home/entrypoint.sh"</span><span class="p">]</span>
|
||||||
|
</code></pre></div>
|
||||||
|
|
||||||
|
<p>And there we have it... an in memory containerised reporting solution with blazing fast capability to aggregate and build reports based on curated data direct from the business.. fully automated and deployable via CI/CD, that provides data updates daily.</p>
|
||||||
|
<p>Now the embedded part.. which isn't built yet but I'll make sure to update you once we have/if we do because the architecture is very exciting for an embbdedded reporting workflow that is deployable via CI/CD processes to applications. As a little taster I'll point you to the <a href="https://www.metabase.com/learn/administration/git-based-workflow">metabase documentation</a>, the unfortunate thing about it is Metabase <em>have</em> hidden this behind the enterprise license.. but I can absolutely see why. If we get to implementing this I'll be sure to update you here on the learnings.</p>
|
||||||
|
<p>Until then....</p>
|
||||||
|
</article>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Footer -->
|
||||||
|
<footer>
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<p>
|
||||||
|
<script type="text/javascript" src="https://sessionize.com/api/speaker/sessions/83c5d14a-bd19-46b4-8335-0ac8358ac46d/0x0x91929ax">
|
||||||
|
</script>
|
||||||
|
</p>
|
||||||
|
<ul class="list-inline text-center">
|
||||||
|
<li>
|
||||||
|
<a href="https://twitter.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-twitter fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://facebook.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-facebook fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://github.com/armistace">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-github fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
<p class="copyright text-muted">Blog powered by <a href="http://getpelican.com">Pelican</a>,
|
||||||
|
which takes great advantage of <a href="http://python.org">Python</a>.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</footer>
|
||||||
|
|
||||||
|
<!-- jQuery -->
|
||||||
|
<script src="http://localhost:8000/theme/js/jquery.js"></script>
|
||||||
|
|
||||||
|
<!-- Bootstrap Core JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/bootstrap.min.js"></script>
|
||||||
|
|
||||||
|
<!-- Custom Theme JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/clean-blog.min.js"></script>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
|
||||||
|
</html>
|
||||||
+8
-59
@@ -11,7 +11,7 @@
|
|||||||
<title>Andrew Ridgway's Blog</title>
|
<title>Andrew Ridgway's Blog</title>
|
||||||
|
|
||||||
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
<link href="http://localhost:8000/feeds/data-engineering.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
<link href="http://localhost:8000/feeds/data-analytics.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
<!-- Bootstrap Core CSS -->
|
<!-- Bootstrap Core CSS -->
|
||||||
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
@@ -38,9 +38,7 @@
|
|||||||
|
|
||||||
|
|
||||||
<meta name="tags" contents="data engineering" />
|
<meta name="tags" contents="data engineering" />
|
||||||
<meta name="tags" contents="DBT" />
|
<meta name="tags" contents="Data Analytics" />
|
||||||
<meta name="tags" contents="Terraform" />
|
|
||||||
<meta name="tags" contents="IAC" />
|
|
||||||
|
|
||||||
|
|
||||||
<meta property="og:locale" content="en">
|
<meta property="og:locale" content="en">
|
||||||
@@ -48,11 +46,11 @@
|
|||||||
|
|
||||||
<meta property="og:type" content="article">
|
<meta property="og:type" content="article">
|
||||||
<meta property="article:author" content="">
|
<meta property="article:author" content="">
|
||||||
<meta property="og:url" content="http://localhost:8000/CI/CD in Data and Data Infrastructure.html">
|
<meta property="og:url" content="http://localhost:8000/notebook-or-bi.html">
|
||||||
<meta property="og:title" content="CI/CD in Data Engineering">
|
<meta property="og:title" content="Notebook or BI, What is the most appropiate communication medium">
|
||||||
<meta property="og:description" content="">
|
<meta property="og:description" content="">
|
||||||
<meta property="og:image" content="http://localhost:8000/">
|
<meta property="og:image" content="http://localhost:8000/">
|
||||||
<meta property="article:published_time" content="2023-06-15 20:00:00+10:00">
|
<meta property="article:published_time" content="2023-07-13 20:00:00+10:00">
|
||||||
</head>
|
</head>
|
||||||
|
|
||||||
<body>
|
<body>
|
||||||
@@ -88,10 +86,10 @@
|
|||||||
<div class="row">
|
<div class="row">
|
||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<div class="post-heading">
|
<div class="post-heading">
|
||||||
<h1>CI/CD in Data Engineering</h1>
|
<h1>Notebook or BI, What is the most appropiate communication medium</h1>
|
||||||
<span class="meta">Posted by
|
<span class="meta">Posted by
|
||||||
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
on Thu 15 June 2023
|
on Thu 13 July 2023
|
||||||
</span>
|
</span>
|
||||||
|
|
||||||
</div>
|
</div>
|
||||||
@@ -106,56 +104,7 @@
|
|||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<!-- Post Content -->
|
<!-- Post Content -->
|
||||||
<article>
|
<article>
|
||||||
<p>Data Engineering has traditionally been considered the bastard step child of work that would have once been considered Administrative in the Tech world. Predominately we write SQL and then deploy that SQL onto one or more Databases. In fact a lot of the traditional methodologies around data almost assume this is the core of how an organistation is managing the majority of it's data. In the last couple of years though there has been a very steady move towards having the Data Engineering workload of SQL move towards Software Engineering techniques. With the popularity of tools like <a href="https://www.dbtlabs.com">DBT</a> and the latest newcommer on the block, <a href="https://www.sqlmesh.com">SQL-MESH</a> The oppportunity has started to arise where we can align our Data Engineering workloads with different environments and move much more efficiently towards a Continous Integration and Deployment methodology in our workflows. </p>
|
<p>I want to preface this post by saying I think "Dashboards" or "BI" as terms are wayyyyyyyyyyyyyyyyy over saturated in the market. There seems to be a belief that any question answerable in data deserves the work associated with a dashboard when in fact a simple one off report, or notebook, would be more than enough.</p>
|
||||||
<p>For the Data Engineering space the move to the cloud has been a breath of fresh air (Not so in some other IT disciplines). I am relatively young, so I don't 100% remember but my experience has taught me that there were 3 options here not so long ago</p>
|
|
||||||
<p><em>Expensive:</em></p>
|
|
||||||
<ul>
|
|
||||||
<li>SAS</li>
|
|
||||||
<li>SSIS/SSRS</li>
|
|
||||||
<li>COGNOS/TM1</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>Rickety:</em></p>
|
|
||||||
<ul>
|
|
||||||
<li>Just write stored procedures!</li>
|
|
||||||
<li>Startup script on my laptop XD</li>
|
|
||||||
<li>"Don't touch that machine over there, No one knows what it does but if it's turned off our financial reports don't work" (This is a third hand story I heard, seriously!)</li>
|
|
||||||
<li>"I need to an upgrade to my laptop, Excel needs more than 8GB of RAM"</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>Hard:</em></p>
|
|
||||||
<ul>
|
|
||||||
<li>Hadoop</li>
|
|
||||||
<li>Spark (hadoop but whatever)</li>
|
|
||||||
<li>Python</li>
|
|
||||||
<li>R</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>(The reason I've listed them as hard is because self hosting Hadoop/Spark and managing a truckload of python or R scripts, whilst it could have been "cheap" required a team of devs who really really knew what they were doing... so not really cheap and also <strong>really hard</strong>)</em></p>
|
|
||||||
<p>Then there was getting git behind all the sql scripts and modelling, let alone CI/CD <strong>IF</strong> it existed, it was custom, and bespoke and likely had a single point of failure in the person who knew how <code>git merge</code> worked. At least... thats I was told I'm not <em>that</em> old ;p.</p>
|
|
||||||
<p>These days we are pretty blessed, with democrotisation of clusters and data Infrastructure in the cloud we no longer need a team of sysadmins who know how to tune a cluster to the Nth degree to get the best our of our data workloads (well... we do, but we pay the cloud guys for that!). However, we still need to know about the idiosyncracities of this infrastructure, when it is appropiate to use and how we want to control and maintain the workloads. </p>
|
|
||||||
<p>In general when I am designing a system I normally like to break it into 3.</p>
|
|
||||||
<ul>
|
|
||||||
<li>Storage</li>
|
|
||||||
<li>Compute</li>
|
|
||||||
<li>Code</li>
|
|
||||||
</ul>
|
|
||||||
<p><em>In General</em> Storage and Compute will be infrastructer related, "Code" is sort of a catch all for my modelling, normally sql, python/r or spark scripts that are used to provide system or business logic, anything really thats going to get data to the end user/analyst/data scientists/annoying person </p>
|
|
||||||
<p>Traditionally the compute layer only really had 2 considerations</p>
|
|
||||||
<ul>
|
|
||||||
<li>SQL or Logic engine (normally a flavour of spark(glue) and then something like reshift/athena/trino/bigquery)</li>
|
|
||||||
<li>Orchestration Layer (Airflow, Dagster)</li>
|
|
||||||
</ul>
|
|
||||||
<p>But with the advent of sql engine agnostic Modelling we potentially now need to also consider</p>
|
|
||||||
<ul>
|
|
||||||
<li>Model Compilation</li>
|
|
||||||
</ul>
|
|
||||||
<p>Now on the surface it seems counterintuitive to seperate the models from the logic layer but lets consider the following scenario</p>
|
|
||||||
<blockquote>
|
|
||||||
<p>Redshift is costing to much and is getting slow, we want to try bigquery
|
|
||||||
How much investment will it be to change over</p>
|
|
||||||
</blockquote>
|
|
||||||
<p>Now, If the entirety of your modelling is stored and deployed to big query direct this would not only involve the investment of spinning up the big query account and either connecting or migrating your data over. You would also need to consider how in the bloody hell you convert all your existing models and workflows over.</p>
|
|
||||||
<p>With something like DBT or sqlmesh you change your compilation target and it's done for you. It also means the Data Engineer now doesn't need to necessarily understand the esoteric nature of the target, at least for simple models (which, lets be real, most are).</p>
|
|
||||||
<p>BUT, now we have <em>a lot</em> of software and infrastructure a simplified common datastack will look something like the below (Assuming ELT, ETL is a bit different but more or less needs the same components)</p>
|
|
||||||
<p><img src="http://localhost:8000/images/DataStackSimplified.png" width="600" height="295" /></p>
|
|
||||||
</article>
|
</article>
|
||||||
|
|
||||||
<hr>
|
<hr>
|
||||||
@@ -0,0 +1,300 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<meta name="description" content="">
|
||||||
|
<meta name="author" content="">
|
||||||
|
|
||||||
|
<title>Andrew Ridgway's Blog</title>
|
||||||
|
|
||||||
|
<link href="http://localhost:8000/feeds/all.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Full Atom Feed" />
|
||||||
|
<link href="http://localhost:8000/feeds/resume.atom.xml" type="application/atom+xml" rel="alternate" title="Andrew Ridgway's Blog Categories Atom Feed" />
|
||||||
|
|
||||||
|
<!-- Bootstrap Core CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/bootstrap.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom CSS -->
|
||||||
|
<link href="http://localhost:8000/theme/css/clean-blog.min.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Code highlight color scheme -->
|
||||||
|
<link href="http://localhost:8000/theme/css/code_blocks/tomorrow.css" rel="stylesheet">
|
||||||
|
|
||||||
|
<!-- Custom Fonts -->
|
||||||
|
<link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
|
||||||
|
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'>
|
||||||
|
|
||||||
|
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
|
||||||
|
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
|
||||||
|
<!--[if lt IE 9]>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
|
||||||
|
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
|
||||||
|
<![endif]-->
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
<meta name="tags" contents="Cover Letter" />
|
||||||
|
<meta name="tags" contents="Resume" />
|
||||||
|
|
||||||
|
|
||||||
|
<meta property="og:locale" content="en">
|
||||||
|
<meta property="og:site_name" content="Andrew Ridgway's Blog">
|
||||||
|
|
||||||
|
<meta property="og:type" content="article">
|
||||||
|
<meta property="article:author" content="">
|
||||||
|
<meta property="og:url" content="http://localhost:8000/resume.html">
|
||||||
|
<meta property="og:title" content="A Resume">
|
||||||
|
<meta property="og:description" content="">
|
||||||
|
<meta property="og:image" content="http://localhost:8000/">
|
||||||
|
<meta property="article:published_time" content="2024-02-23 20:00:00+10:00">
|
||||||
|
</head>
|
||||||
|
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<!-- Navigation -->
|
||||||
|
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
|
||||||
|
<div class="container-fluid">
|
||||||
|
<!-- Brand and toggle get grouped for better mobile display -->
|
||||||
|
<div class="navbar-header page-scroll">
|
||||||
|
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
|
||||||
|
<span class="sr-only">Toggle navigation</span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
<span class="icon-bar"></span>
|
||||||
|
</button>
|
||||||
|
<a class="navbar-brand" href="http://localhost:8000/">Andrew Ridgway's Blog</a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Collect the nav links, forms, and other content for toggling -->
|
||||||
|
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
|
||||||
|
<ul class="nav navbar-nav navbar-right">
|
||||||
|
|
||||||
|
</ul>
|
||||||
|
</div>
|
||||||
|
<!-- /.navbar-collapse -->
|
||||||
|
</div>
|
||||||
|
<!-- /.container -->
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<!-- Page Header -->
|
||||||
|
<header class="intro-header" style="background-image: url('http://localhost:8000/theme/images/post-bg.jpg')">
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<div class="post-heading">
|
||||||
|
<h1>A Resume</h1>
|
||||||
|
<span class="meta">Posted by
|
||||||
|
<a href="http://localhost:8000/author/andrew-ridgway.html">Andrew Ridgway</a>
|
||||||
|
on Fri 23 February 2024
|
||||||
|
</span>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<!-- Main Content -->
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<!-- Post Content -->
|
||||||
|
<article>
|
||||||
|
<h1>OVERVIEW</h1>
|
||||||
|
<p>I am a Senior Data Engineer looking to transition my skills to Data and Solution
|
||||||
|
Architecting as well as project management. I have spent the better part of the
|
||||||
|
last decade refining my abilities in taking business requirements and turning
|
||||||
|
those into actionable data engineering, analytics, and software projects with
|
||||||
|
trackable metrics. I believe in agnosticism when it comes to coding languages
|
||||||
|
and have experimented in my own time with many different languages. In my
|
||||||
|
career I have used Python, .NET, PowerShell, TSQL, VB and SAS (multiple
|
||||||
|
products) in an Enterprise capacity. I also have experience using Google Cloud
|
||||||
|
Platform and AWS tools for ETL and data platform development as well as git
|
||||||
|
for version control and deployment using various IAC tools. I have also
|
||||||
|
conducted data analysis and modelling on business metrics to find relationships
|
||||||
|
between both staff and customer behavior and produced actionable
|
||||||
|
recommendations based on the conclusions. In a private context I have also
|
||||||
|
experimented with C, C# and Kotlin I am looking to further my career by taking
|
||||||
|
my passion for data engineering and analysis as well as web and software
|
||||||
|
development and applying it in a strategic context.</p>
|
||||||
|
<h1>SKILLS & ABILITIES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Python (scripting, compiling, notebooks – Sagemaker, Jupyter)</li>
|
||||||
|
<li>git</li>
|
||||||
|
<li>SAS (Base, EG, VA)</li>
|
||||||
|
<li>Various Google Cloud Tools (Data Fusion, Compute Engine, Cloud Functions)</li>
|
||||||
|
<li>Various Amazon Tools (EC2, RDS, Kinesis, Glue, Redshift, Lambda, ECS, ECR, EKS)</li>
|
||||||
|
<li>Streaming Technologies (Kafka, Hive, Spark Streaming)</li>
|
||||||
|
<li>Various DB platforms both on Prem and Serverless (MariaDB/MySql,</li>
|
||||||
|
<li>Postgres/Redshift, SQL Server, RDS/Aurora variants)</li>
|
||||||
|
<li>Various Microsoft Products (PowerBI, TSQL, Excel, VBA)</li>
|
||||||
|
<li>Linux Server Administration (cron, bash, systemD)</li>
|
||||||
|
<li>ETL/ELT Development</li>
|
||||||
|
<li>Basic Data Modelling (Kimball, SCD Type 2)</li>
|
||||||
|
<li>IAC (Cloud Formation, Terraform)</li>
|
||||||
|
<li>Datahub Deployment</li>
|
||||||
|
<li>Dagster Orchestration Deployments</li>
|
||||||
|
<li>DBT Modelling and Design Deployments</li>
|
||||||
|
<li>Containerised and Cloud Driven Data Architecture</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EXPERIENCE</h1>
|
||||||
|
<h2>Cloud Data Architect</h2>
|
||||||
|
<h3><em>Redeye Apps</em></h3>
|
||||||
|
<h4><em>May 2022 - Present</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Greenfields Research, Design and Deployment of S3 datalake (Parquet)</li>
|
||||||
|
<li>AWS DMS, S3, Athena, Glue</li>
|
||||||
|
<li>Research Design and Deployment of Catalog (Datahub)</li>
|
||||||
|
<li>Design of Data Governance Process (Datahub driven)</li>
|
||||||
|
<li>Research Design and Deployment of Orchestration and Modelling for Transforms (Dagster/DBT into Mesos)</li>
|
||||||
|
<li>CI/CD design and deployment of modelling and orchestration using Gitlab</li>
|
||||||
|
<li>Research, Design and Deployment of ML Ops Dev pipelines anddeployment strategy</li>
|
||||||
|
<li>Design of ETL/Pipelines (DBT)</li>
|
||||||
|
<li>Design of Customer Facing Data Products and deployment methodologies (Fully automated via Kakfa/Dagster/DBT)</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Data Engineer,</h2>
|
||||||
|
<h3><em>TechConnect IT Solutions</em></h3>
|
||||||
|
<h4><em>August 2021 – May 2022</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Design of Cloud Data Batch ETL solutions using Python (Glue)</li>
|
||||||
|
<li>Design of Cloud Data Streaming ETL solution using Python (Kinesis)</li>
|
||||||
|
<li>Solve complex client business problems using software to join and transform data from DB’s, Web API’s, Application API’s and System logs</li>
|
||||||
|
<li>Build CI/CD pipelines to ensure smooth deployments (Bitbucket, gitlab)</li>
|
||||||
|
<li>Apply Prebuilt ML models to software solutions (Sagemaker)</li>
|
||||||
|
<li>Assist with the architecting of Containerisation solutions (Docker, ECS, ECR)</li>
|
||||||
|
<li>API testing and development (gRPC, Rest)</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Enterprise Data Warehouse Developer</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>August 2019 - August 2021</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>ETL development of CRM, WFP, Outbound Dialer, Inbound switch in Google Cloud, SAS, TSQL</li>
|
||||||
|
<li>Bringing new data to the business to analyse for new insights</li>
|
||||||
|
<li>Redeveloped Version Control and brought git to the data team</li>
|
||||||
|
<li>Introduced python for API enablement in the Enterprise Data Warehouse</li>
|
||||||
|
<li>Partnering with the business to focus data project on actual need and translating into technical requirements</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Business Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2018 - August 2019</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Automate Service Performance Reporting using PowerShell/VBA/SAS</li>
|
||||||
|
<li>Learn and leverage SAS EG and VA to streamline Microsoft Excel Reporting</li>
|
||||||
|
<li>Identify and develop data pipelines to source data from multiple sources easily and collate into a single source to identify relationships and trends</li>
|
||||||
|
<li>Technologies used include VBA, PowerShell, SQL, Web API’s, SAS</li>
|
||||||
|
<li>Where SAS is inappropriate use VBA to automate processes in Microsoft Access and Excel</li>
|
||||||
|
<li>Gather Requirements to build meaningful reporting solutions</li>
|
||||||
|
<li>Provide meaningful analysis on business performance and provide relevant presentations and reports to senior stakeholders.</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Forecasting and Capacity Analyst</h2>
|
||||||
|
<h3><em>Auto and General Insurance</em></h3>
|
||||||
|
<h4><em>January 2017 – January 2018</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop the outbound forecasting model for the Auto and General sales call center by analysing the relationship between customer decisions and workload drivers</li>
|
||||||
|
<li>This includes the complete data pipeline for the model from identifying and sourcing data, building the reporting and analysing the data and associated drivers.</li>
|
||||||
|
<li>Forecast inbound workload requirements for the Auto and General sales call center using time series analysis</li>
|
||||||
|
<li>Learn and leverage the Aspect Workforce Management System to ensure efficiency of forecast generation</li>
|
||||||
|
<li>Learn and leverage the capabilities of SAS Enterprise Guide to improve accuracy</li>
|
||||||
|
<li>Liaise with people across the business to ensure meaningful, accurate analysis is provided to senior stakeholders</li>
|
||||||
|
<li>Analyse monthly, weekly and intraday requirements and ensure forecast is accurately predicting workload for breaks, meetings and Leave</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Senior HR Performance Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>June 2016 - January 2017</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Harmonise various systems to develop a unified workforce reporting and analysis framework with appropriate metrics</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h2>Workforce Business Analyst</h2>
|
||||||
|
<h3><em>Queensland Department of Justice and Attorney General</em></h3>
|
||||||
|
<h4><em>July 2015 – June 2016</em></h4>
|
||||||
|
<ul>
|
||||||
|
<li>Develop and refine current workforce analysis techniques and databases</li>
|
||||||
|
<li>Use VBA to automate regular reporting in Microsoft Access and Excel</li>
|
||||||
|
<li>Act as liaison between shared service providers and executives and facilitate communication during the implementation of a payroll leave audit</li>
|
||||||
|
<li>Gather reporting requirements from various business areas and produce ad-hoc and regular reports as required</li>
|
||||||
|
<li>Participate in government process through the production of briefs including Questions on Notice and Estimates Briefs for departmental executives</li>
|
||||||
|
</ul>
|
||||||
|
<h1>EDUCATION</h1>
|
||||||
|
<ul>
|
||||||
|
<li>2011 Bachelor of Business Management, University of Queensland</li>
|
||||||
|
<li>2008 Bachelor of Arts, University of Queensland</li>
|
||||||
|
</ul>
|
||||||
|
<h1>REFERENCES</h1>
|
||||||
|
<ul>
|
||||||
|
<li>Anthony Stiller Lead Developer, Data warehousing, Queensland Health</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0428 038 031</em></p>
|
||||||
|
<ul>
|
||||||
|
<li>Jaime Brian Head of Cloud Ninjas, TechConnect</li>
|
||||||
|
</ul>
|
||||||
|
<p><em>0422 012 17</em></p>
|
||||||
|
</article>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<hr>
|
||||||
|
|
||||||
|
<!-- Footer -->
|
||||||
|
<footer>
|
||||||
|
<div class="container">
|
||||||
|
<div class="row">
|
||||||
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
|
<p>
|
||||||
|
<script type="text/javascript" src="https://sessionize.com/api/speaker/sessions/83c5d14a-bd19-46b4-8335-0ac8358ac46d/0x0x91929ax">
|
||||||
|
</script>
|
||||||
|
</p>
|
||||||
|
<ul class="list-inline text-center">
|
||||||
|
<li>
|
||||||
|
<a href="https://twitter.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-twitter fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://facebook.com/ar17787">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-facebook fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
<li>
|
||||||
|
<a href="https://github.com/armistace">
|
||||||
|
<span class="fa-stack fa-lg">
|
||||||
|
<i class="fa fa-circle fa-stack-2x"></i>
|
||||||
|
<i class="fa fa-github fa-stack-1x fa-inverse"></i>
|
||||||
|
</span>
|
||||||
|
</a>
|
||||||
|
</li>
|
||||||
|
</ul>
|
||||||
|
<p class="copyright text-muted">Blog powered by <a href="http://getpelican.com">Pelican</a>,
|
||||||
|
which takes great advantage of <a href="http://python.org">Python</a>.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</footer>
|
||||||
|
|
||||||
|
<!-- jQuery -->
|
||||||
|
<script src="http://localhost:8000/theme/js/jquery.js"></script>
|
||||||
|
|
||||||
|
<!-- Bootstrap Core JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/bootstrap.min.js"></script>
|
||||||
|
|
||||||
|
<!-- Custom Theme JavaScript -->
|
||||||
|
<script src="http://localhost:8000/theme/js/clean-blog.min.js"></script>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
|
||||||
|
</html>
|
||||||
@@ -83,11 +83,13 @@
|
|||||||
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
|
||||||
<h1>Tags for Andrew Ridgway's Blog</h1> <li><a href="http://localhost:8000/tag/amazon.html">Amazon</a> (1)</li>
|
<h1>Tags for Andrew Ridgway's Blog</h1> <li><a href="http://localhost:8000/tag/amazon.html">Amazon</a> (1)</li>
|
||||||
<li><a href="http://localhost:8000/tag/containers.html">containers</a> (1)</li>
|
<li><a href="http://localhost:8000/tag/containers.html">containers</a> (1)</li>
|
||||||
|
<li><a href="http://localhost:8000/tag/cover-letter.html">Cover Letter</a> (2)</li>
|
||||||
<li><a href="http://localhost:8000/tag/data-engineering.html">data engineering</a> (3)</li>
|
<li><a href="http://localhost:8000/tag/data-engineering.html">data engineering</a> (3)</li>
|
||||||
<li><a href="http://localhost:8000/tag/dbt.html">DBT</a> (1)</li>
|
<li><a href="http://localhost:8000/tag/duckdb.html">DuckDB</a> (1)</li>
|
||||||
<li><a href="http://localhost:8000/tag/iac.html">IAC</a> (1)</li>
|
<li><a href="http://localhost:8000/tag/embedded.html">embedded</a> (1)</li>
|
||||||
<li><a href="http://localhost:8000/tag/managed-services.html">Managed Services</a> (1)</li>
|
<li><a href="http://localhost:8000/tag/managed-services.html">Managed Services</a> (1)</li>
|
||||||
<li><a href="http://localhost:8000/tag/terraform.html">Terraform</a> (1)</li>
|
<li><a href="http://localhost:8000/tag/metabase.html">Metabase</a> (1)</li>
|
||||||
|
<li><a href="http://localhost:8000/tag/resume.html">Resume</a> (2)</li>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
|
|||||||
Reference in New Issue
Block a user