Skip to content

Latest commit

 

History

303 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

schoolmint-connector

Data pipeline for Schoolmint application data. This pipeline pings Schoolmint's API to generate a report that is sent to an SFTP server. The pipeline then loads that file from the SFTP server to Google Cloud Storage.

Dependencies:

Additional Dependencies

This job requires an FTP server for Schoolmint to drop the files on. Our org uses a Digital Ocean droplet, but a SFTP server through any service should work. You will need to provide some info about the server to Schoolmint so they can create your API token.

Getting Started

Setup Environment

  1. Clone this repo
$ git clone https://github.com/kippnorcal/ukg-connector.git 
  1. Install Pipenv
$ pip install pipenv
$ pipenv install
  1. Install Docker
  1. Create .env file with project secrets
DELETE_LOCAL_FILES=1
REJECT_EMPTY_FILES=0

# Google Storage Info
GOOGLE_APPLICATION_CREDENTIALS= # Path to your cred file
BUCKET= # Name of your Cloud Storage Bucket

# FTP Credential Variables
FTP_HOSTNAME=
FTP_USERNAME=
FTP_PWD=
FTP_KEY=
ARCHIVE_MAX_DAYS=60

# API Variables
API_SUFFIXES=
API_DOMAIN_REGIONAL=
API_ACCOUNT_EMAIL_REGIONAL=
API_TOKEN_REGIONAL=

# Slack Notifications
FROM_ADDRESS=
TO_ADDRESS=
MG_API_KEY=
MG_API_URL=
MG_DOMAIN=
FAILURE_EMAIL=  # Address of Slack channel for failure emails

# dbt variables
DBT_ACCOUNT_ID=
DBT_JOB_ID=
DBT_BASE_URL=
DBT_PERSONAL_ACCESS_TOKEN=

Note on API_SUFFIXES

Our org had a need to get Schoolmint data from multiple API's that had the same name except for the final word in the name. The automation will iterate over each suffix listed in API_SUFFIXES and join them with API_DOMAIN_REGIONAL to get the full endpoint name. In the future, this will be removed since we are no long using multiple endpoints.

  1. Build Docker Image
$ docker build -t schoolmint .

Running the Job

Here are the runtime arguments for the job:

Arg Description
--school-year Required; Determines which school year the job is processing; The value of this arg gets added to a school_year_4_digit field
--dbt-refresh Optional; Can run a job in dbt to refresh Schoolmint data
--semt-refresh Optional; Runs an adhoc dbt job to refresh the source of the SEMT trackers
-a, --add-cols-hist Optional; When columns are added or removed to the current year's report, the historical reports need to be updated as well. This command will add or remove columns based on the current year's fields.
-g, --generate-schema-json Optional; When columns are added or removed from the report, the external table in BigQuery needs to be recreated with a new schema. This command creates a json version of the schema that can be copied and pasted in. This command requires a volumne to be mount to the docker container using -v

Examples

Run the job:

docker run --rm -t schoolmint --scool-year 2027

Update the columns of historical files:

docker run --rm -t schoolmint -a --scool-year 2027

Generate a JSON file for a new schema (using the -v volume flag):

docker run --rm -t -v ~/Desktop:/code/files schoolmint -g --school-year 2027

Maintenance

The token for Schoolmint's API needs to be updated annually. This can be updated in the .env file. Making changes to the report through Schoolmint's UI will break the automation. Changes require a new API token.

About

Load all SchoolMint application and lottery data into database for analysis.

Topics

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages