<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>by.Maksim</title>
    <link>https://bymaksim.com/</link>
    <description>Government | Technology | Data | Maps</description>
    <pubDate>Sun, 27 Sep 2026 03:03:20 +0000</pubDate>
    <item>
      <title>COVID-19 Mobility Reports from Google</title>
      <link>https://bymaksim.com/covid-19-mobility-reports-from-google?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[#covid19 #data #government&#xA;&#xA;Google released a set of community mobility reports using their anonymized movement data from Google Maps:&#xA;&#xA;San Diego Mobility Data&#xA;&#xA;!--more--&#xA;&#xA;Officials could use these to compare policies and communication tactics across similar jurisdictions. I do wish the data were more granular, but I understand the privacy implications that would pose. &#xA;&#xA;Shaming South Dakota&#xA;Let&#39;s compare Sioux Falls South Dakota with Louisville, Kentucky - two similar cities.&#xA;&#xA;As of March 31, South Dakota is the only state not implementing social distancing or shelter-in-place measures. Kentucky is.&#xA;&#xA;CovidActNow&#xA;&#xA;Here is the mobility report for Minnehaha County that contains Sioux Falls. Sioux Falls is South Dakota&#39;s most populous city.&#xA;&#xA;Sioux Falls Mobility&#xA;&#xA;Compare that with Jefferson County, containing Louisville, KY:&#xA;&#xA;Kentucky Mobility &#xA;&#xA;There&#39;s a more substantial drop in retail and workplace attendance in Louisville in Sioux Falls. There is also a higher increase in park attendance.&#xA;&#xA;There&#39;s a lot of nuances, though.  For example, Sioux Falls has COVID19 guidance their site.  But without a statewide order, it may not be striking a lot of the population as necessary.  &#xA;&#xA;Californians hate parks?&#xA;I also found this interesting.  Louisville and Sioux Falls had a significant spike in park attendance. San Diego County had a significant drop.  &#xA;&#xA;San Diego County Park Drop&#xA;&#xA;It&#39;s possible San Diegans hate parks. Or - more likely -  the end of winter in Kentucky and South Dakota was encouraging people to go outside.&#xA;&#xA;Downloading the Reports&#xA;&#xA;To download these, head over to the Community Mobility Reports page, and find your state. &#xA;&#xA;The first page of the report shows aggregated metrics for the state, with an explanation of what they are:&#xA;&#xA;Google Mobility Report First Page&#xA;&#xA;As you scroll down, you can get individual county reports.&#xA;&#xA;If you end up finding them useful, please let me know how by tweeting me @MrMaksimize.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p><a href="https://bymaksim.com/tag:covid19" class="hashtag" rel="nofollow"><span>#</span><span class="p-category">covid19</span></a> <a href="https://bymaksim.com/tag:data" class="hashtag" rel="nofollow"><span>#</span><span class="p-category">data</span></a> <a href="https://bymaksim.com/tag:government" class="hashtag" rel="nofollow"><span>#</span><span class="p-category">government</span></a></p>

<p>Google released a set of <a href="https://www.google.com/covid19/mobility" rel="nofollow">community mobility reports</a> using their anonymized movement data from Google Maps:</p>

<p><img src="https://take.ms/iNBtL" alt="San Diego Mobility Data"/></p>



<p>Officials could use these to compare policies and communication tactics across similar jurisdictions. I do wish the data were more granular, but I understand the privacy implications that would pose.</p>

<h2 id="shaming-south-dakota">Shaming South Dakota</h2>

<p>Let&#39;s compare Sioux Falls South Dakota with Louisville, Kentucky – two similar cities.</p>

<p>As of March 31, South Dakota <a href="https://covidactnow.org" rel="nofollow">is the only state not implementing social distancing or shelter-in-place measures</a>. Kentucky is.</p>

<p><img src="https://take.ms/yiQRV" alt="CovidActNow"/></p>

<p>Here is the mobility report for Minnehaha County that contains Sioux Falls. Sioux Falls is South Dakota&#39;s most populous city.</p>

<p><img src="https://take.ms/oTNSE" alt="Sioux Falls Mobility"/></p>

<p>Compare that with Jefferson County, containing Louisville, KY:</p>

<p><img src="https://take.ms/mySgtA" alt="Kentucky Mobility"/></p>

<p>There&#39;s a more substantial drop in retail and workplace attendance in Louisville in Sioux Falls. There is also a higher increase in park attendance.</p>

<p>There&#39;s a lot of nuances, though.  For example, Sioux Falls has COVID19 <a href="https://www.siouxfalls.org/covid19" rel="nofollow">guidance</a> their site.  But without a statewide order, it may not be striking a lot of the population as necessary.</p>

<h2 id="californians-hate-parks">Californians hate parks?</h2>

<p>I also found this interesting.  Louisville and Sioux Falls had a significant spike in park attendance. San Diego County had a significant drop.</p>

<p><img src="https://take.ms/krIev" alt="San Diego County Park Drop"/></p>

<p>It&#39;s possible San Diegans hate parks. Or – more likely –  the end of winter in Kentucky and South Dakota was encouraging people to go outside.</p>

<h2 id="downloading-the-reports">Downloading the Reports</h2>

<p>To download these, head over to the <a href="https://www.google.com/covid19/mobility/" rel="nofollow">Community Mobility Reports</a> page, and find your state.</p>

<p>The first page of the report shows aggregated metrics for the state, with an explanation of what they are:</p>

<p><img src="https://take.ms/4RAQP" alt="Google Mobility Report First Page"/></p>

<p>As you scroll down, you can get individual county reports.</p>

<p>If you end up finding them useful, please let me know how by tweeting me <a href="https://twitter.com/MrMaksimize" rel="nofollow">@MrMaksimize</a>.</p>
]]></content:encoded>
      <guid>https://bymaksim.com/covid-19-mobility-reports-from-google</guid>
      <pubDate>Fri, 03 Apr 2020 12:10:49 +0000</pubDate>
    </item>
    <item>
      <title>Telling a good story is not the priority.</title>
      <link>https://bymaksim.com/telling-a-good-story-is-not-the-priority?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[It’s the side effect.&#xA;!--more--&#xA;Photo by rawpixel on [Unsplash](https://cdn-images-1.medium.com/max/5000/0VQmXb0JVtEIaE1-5)Photo by rawpixel on Unsplash*&#xA;&#xA;Successfully applying data and technology in government is hard work. Telling a good story about these applications is just as challenging.&#xA;&#xA;“What makes a good data story?” does not interest me.&#xA;&#xA;“What makes a successful data project?” interests me.&#xA;“How do successful data teams operate?” interests me.&#xA;“How do we perpetually sustain good government?” interests me.&#xA;&#xA;A good data story&#xA;&#xA;A good story doesn’t care about the tech or the tools.&#xA;It cares about the people that “led the effort”;&#xA;The politician that got credit for it;&#xA;Or the organization that wrote about it.&#xA;&#xA;A good data story has just the right amount of buzzwords.&#xA;It will fit into a 1,000 word blog post (like this one),&#xA;Regardless of the underlying project’s value or effectiveness.&#xA;&#xA;A good data story is about using machine learning,&#xA;In the form of advanced neural networks,&#xA;Across data gathered by a $5 million innovative sensor array,&#xA;That stores the data in a Hadoop cluster,&#xA;To determine how many drones just flew by.&#xA;&#xA;Yet, nobody asked how many drones just flew by.&#xA;&#xA;A good data story chases talking points — and forgets to solve a problem that matters.&#xA;&#xA;Sorry, that’s not interesting.&#xA;&#xA;A good data project&#xA;&#xA;A good data project identifies a problem that is worth solving.&#xA;It requires that city staff across departments work together;&#xA;Means building buy-in for data access and process change;&#xA;And accounts for regulations that are decades old.&#xA;&#xA;A good data project is about details,&#xA;The right technical approach to create the right solution,&#xA;And the process change that influences the right outcome.&#xA;&#xA;A good data project is about the Streets engineers who spent years understanding street paving,&#xA;So they can spend 20 hours explaining their processes to the data scientists,&#xA;Who will to figure out how to run Python on our legacy servers,&#xA;So that the front-end developers can provide a usable interface to street paving analysis,&#xA;And the program manager who championed the project can implement it operationally.&#xA;&#xA;A good data team&#xA;&#xA;A good data team builds relationships with employees across all city disciplines;&#xA;They use design thinking techniques to identify hard problems;&#xA;And make disciplined technology choices to deliver the right solution.&#xA;&#xA;They save countless hours for Streets engineers,&#xA;By helping them them auto-generate reports.&#xA;&#xA;They analyze historical data and inform dispatch configurations for Fire/EMS,&#xA;Helping doctors make the case for adjusting dispatch plans.&#xA;&#xA;They evaluate fleet vehicle GPS movement and propose new routing patterns for delivery trucks,&#xA;Allowing for more efficient movement of materials between city offices.&#xA;&#xA;They build Police dispatchers a real-time dashboard with emergency call hold times,&#xA;By installing a Firefox plugin.&#xA;&#xA;A good government&#xA;&#xA;The aforementioned projects are impactful;&#xA;Cost almost nothing;&#xA;And will likely never see a press release.&#xA;&#xA;We’re okay with that.&#xA;&#xA;We will tell our story to the people that benefited from our work:&#xA;The street engineer who will now save time running reports.&#xA;The truck driver that spends less time in traffic, completing deliveries faster.&#xA;The dispatcher who will more effectively take calls, saving more lives.&#xA;&#xA;We are not here to tell good data stories.&#xA;We are here to do good work.&#xA;We are here for good government.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>It’s the side effect.

<img src="https://cdn-images-1.medium.com/max/5000/0*VQmXb0JVtEIaE1-5" alt="Photo by [rawpixel](https://unsplash.com/@rawpixel?utm_source=medium&amp;utm_medium=referral) on [Unsplash](https://unsplash.com?utm_source=medium&amp;utm_medium=referral)"/><em>Photo by <a href="https://unsplash.com/@rawpixel?utm_source=medium&amp;utm_medium=referral" rel="nofollow">rawpixel</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral" rel="nofollow">Unsplash</a></em></p>

<p>Successfully applying data and technology in government is hard work. Telling a good story about these applications is just as challenging.</p>

<p>“What makes a good data story?” does not interest me.</p>

<p>“What makes a successful data project?” interests me.
“How do successful data teams operate?” interests me.
“How do we perpetually sustain good government?” interests me.</p>

<h2 id="a-good-data-story"><strong>A good data story</strong></h2>

<p>A good story doesn’t care about the tech or the tools.
It cares about the people that “led the effort”;
The politician that got credit for it;
Or the organization that wrote about it.</p>

<p>A good data story has just the right amount of buzzwords.
It will fit into a 1,000 word blog post (like this one),
Regardless of the underlying project’s value or effectiveness.</p>

<p>A good data story is about using machine learning,
In the form of advanced neural networks,
Across data gathered by a $5 million innovative sensor array,
That stores the data in a Hadoop cluster,
To determine how many drones just flew by.</p>

<p>Yet, nobody asked how many drones just flew by.</p>

<p>A good data story chases talking points — and forgets to solve a problem that matters.</p>

<p>Sorry, that’s not interesting.</p>

<h2 id="a-good-data-project"><strong>A good data project</strong></h2>

<p>A good data project identifies a problem that is worth solving.
It requires that city staff across departments work together;
Means building buy-in for data access and process change;
And accounts for regulations that are decades old.</p>

<p>A good data project is about details,
The right technical approach to create the right solution,
And the process change that influences the right outcome.</p>

<p>A good data project is about the Streets engineers who spent years understanding street paving,
So they can spend 20 hours explaining their processes to the data scientists,
Who will to figure out how to run Python on our legacy servers,
So that the front-end developers can provide a usable interface to street paving analysis,
And the program manager who championed the project can implement it operationally.</p>

<h2 id="a-good-data-team"><strong>A good data team</strong></h2>

<p>A good data team builds relationships with employees across all city disciplines;
They use design thinking techniques to identify hard problems;
And make disciplined technology choices to deliver the right solution.</p>

<p>They save countless hours for Streets engineers,
By helping them them auto-generate reports.</p>

<p>They analyze historical data and inform dispatch configurations for Fire/EMS,
Helping doctors make the case for adjusting dispatch plans.</p>

<p>They evaluate fleet vehicle GPS movement and propose new routing patterns for delivery trucks,
Allowing for more efficient movement of materials between city offices.</p>

<p>They build Police dispatchers a real-time dashboard with emergency call hold times,
By installing a Firefox plugin.</p>

<h2 id="a-good-government"><strong>A good government</strong></h2>

<p>The aforementioned projects are impactful;
Cost almost nothing;
And will likely never see a press release.</p>

<p>We’re okay with that.</p>

<p>We will tell our story to the people that benefited from our work:
The street engineer who will now save time running reports.
The truck driver that spends less time in traffic, completing deliveries faster.
The dispatcher who will more effectively take calls, saving more lives.</p>

<p>We are not here to tell good data stories.
We are here to do good work.
We are here for good government.</p>
]]></content:encoded>
      <guid>https://bymaksim.com/telling-a-good-story-is-not-the-priority</guid>
      <pubDate>Tue, 19 Jun 2018 00:33:40 +0000</pubDate>
    </item>
    <item>
      <title>A few updates to StreetsSD</title>
      <link>https://bymaksim.com/a-few-updates-to-streetssd?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[Hey there everyone! For all the fans of StreetsSD we have some great news!&#xA;&#xA;We recently completed a small sprint and made a few fixes and improvements.&#xA;&#xA;!--more--&#xA;We love listening to our users, and one of the most requested features has been a way to type in an address and see what streets have been or will be repaired around where you live or work.&#xA;&#xA;We implemented an autocompleting search using Mapzen’s awesome search api. Now you can search for a place (like Balboa Park), or an address and it will zoom right in for you:&#xA;&#xA;We also made a few performance tweaks by moving hosting from Github Pages to Amazon S3 and CloudFront to align with our overall infrastructure and deployment strategy. Hopefully sometime soon we’ll be able to make a few tweaks like we did on data.sandiego.gov to significantly increase the loading speed.&#xA;&#xA;Finally, there is one more tweak we made, but we’re not ready to talk about it yet, because it needs a bigger piece before it goes into production. However, you can probably find it by looking at the open source code driving StreetsSD.&#xA;&#xA;As always, we learn, iterate and continuously improve our products. We Love To Hear and Implement your suggestions!!&#xA;&#xA;But we love pull requests even more!&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>Hey there everyone! For all the fans of StreetsSD we have some great news!</p>

<p>We recently completed a small sprint and made a few fixes and improvements.</p>



<p>We love listening to our users, and one of the most requested features has been a way to type in an address and see what streets have been or will be repaired around where you live or work.</p>

<p>We implemented an autocompleting search using Mapzen’s awesome <a href="https://mapzen.com/products/search/" rel="nofollow">search api</a>. Now you can search for a place (like Balboa Park), or an address and it will zoom right in for you:</p>

<p><img src="https://cdn-images-1.medium.com/max/2558/0*huoKM8MflBxNOg33.jpg" alt=""/></p>

<p>We also made a few performance tweaks by moving hosting from Github Pages to Amazon S3 and CloudFront to align with our overall infrastructure and deployment strategy. Hopefully sometime soon we’ll be able to make a few tweaks <a href="https://data.sandiego.gov/stories/portal-speedup/" rel="nofollow">like we did on data.sandiego.gov</a> to significantly increase the loading speed.</p>

<p>Finally, there is one more tweak we made, but we’re not ready to talk about it yet, because it needs a bigger piece before it goes into production. However, you can probably find it by <a href="https://github.com/cityofsandiego/streetsSD" rel="nofollow">looking at the open source code</a> driving StreetsSD.</p>

<p>As always, we learn, iterate and continuously improve our products. <a href="https://github.com/cityofsandiego/streetsSD/issues" rel="nofollow">We Love To Hear and Implement your suggestions!</a>!</p>

<p>But we love pull requests even more!</p>
]]></content:encoded>
      <guid>https://bymaksim.com/a-few-updates-to-streetssd</guid>
      <pubDate>Wed, 14 Mar 2018 01:35:41 +0000</pubDate>
    </item>
    <item>
      <title>Chief Data and Sustainability Officers form an effective team.</title>
      <link>https://bymaksim.com/chief-data-and-sustainability-officers-form-an-effective-team?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[The position of a Chief Data Officer is a recent position in many North American Cities. I’m not going to dwell on it too much, as there are many articles and talks about this online. In San Diego, I run the Data &amp; Analytics team. We are responsible for Open Data, data integration and automation, and various data projects. Every organization is different though.&#xA;&#xA;The Chief Sustainability Officer is a recent position that has been popping up as well. The CSO works to plan for upcoming challenges brought on by a changing climate. She also works on reducing the City’s greenhouse gas emissions and carbon footprint.&#xA;&#xA;!--more--&#xA;&#xA;Both roles involve dealing with City and regional data. Cities have a trove of data they collect, store and manage, but it’s usually not managed in a strategic way. The CDO knows where the data is, what the caveats are, and how to connect to it and perform automated analysis. The CSO needs continued information on policy baselines and to evaluate policy decisions. Neither of us want to continue to receive data via email or in a PDF.&#xA;&#xA;Both roles have very little definition, very little staff and almost no resources. The CSO and the CDO, have to run their teams as lean startups. Multi-million and multi-year IT projects are off the table. Both are outside-the-box thinkers and consider “that’s the way we’ve always done things” as a challenge.&#xA;&#xA;At the City of San Diego, the CDO and the CSO are natural partners. The CSO has a deep understanding of policy and business objectives. Technology and data, are not always her cup of tea.&#xA;&#xA;The CDO has a strong technical background. He knows where the data is, what it means, and what the caveats are. He knows how to leverage it solve business objectives and watch policy decisions. He doesn’t want to sit in meetings, trying to form policy juggling people’s egos. That’s where the CSO comes in.&#xA;&#xA;Let’s take a concrete use case. As I’m sure you’ve heard, the City of San Diego is getting a deployment of Smart Street streetlights. It’s the first and largest Internet of Things deployment in a city in North America. &#xA;These streetlight sensors have a lot of different capabilities (blog post coming soon). They have the potential to improve parking management. City staff can use them to monitor pedestrian and car traffic. We can use them to track our environment.&#xA;&#xA;Without reading the API documentation it’s hard to understand what’s possible. It takes someone with a background in software to do that. Then, he would have to translate it to humanese. The CDO happens to enjoy reading that kind of stuff. The CSO definitely does not.&#xA;&#xA;The CDO has a full understanding of the sensors’ capabilities. The CSO has in-depth knowledge of what we need to monitor and change to meet climate goals. Together, we can decide how to utilize this technology to meet her business objectives.&#xA;&#xA;These sensors also need a coordinated policy effort to manage the data. The CSO happens to be good at that as policymaking is part of her job. She can help me push through the policy work that I need.&#xA;&#xA;Our unique skills and expertise complement our shared drive for change and our similar desire to get sht done.&#xA;&#xA;The data movement in cities has followed an upward curve similar to the sustainability movement. They have also followed a similar time frame. This is not by accident — both of these movements are new ways of thinking and new policy directions. This is not simple correlation. These roles are very much intertwined and are supporting and enabling each other’s day to day work.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>The position of a Chief Data Officer is a recent position in many North American Cities. I’m not going to dwell on it too much, as there are many articles and talks about this online. In San Diego, I run the Data &amp; Analytics team. We are responsible for Open Data, data integration and automation, and various data projects. Every organization is different though.</p>

<p>The Chief Sustainability Officer is a recent position that has been popping up as well. The CSO works to plan for upcoming challenges brought on by a changing climate. She also works on reducing the City’s greenhouse gas emissions and carbon footprint.</p>



<p>Both roles involve dealing with City and regional data. Cities have a trove of data they collect, store and manage, but it’s usually not managed in a strategic way. The CDO knows where the data is, what the caveats are, and how to connect to it and perform automated analysis. The CSO needs continued information on policy baselines and to evaluate policy decisions. Neither of us want to continue to receive data via email or in a PDF.</p>

<p>Both roles have very little definition, very little staff and almost no resources. The CSO and the CDO, have to run their teams as lean startups. Multi-million and multi-year IT projects are off the table. Both are outside-the-box thinkers and consider “that’s the way we’ve always done things” as a challenge.</p>

<p>At the City of San Diego, the CDO and the CSO are natural partners. The CSO has a deep understanding of policy and business objectives. Technology and data, are not always her cup of tea.</p>

<p>The CDO has a strong technical background. He knows where the data is, what it means, and what the caveats are. He knows how to leverage it solve business objectives and watch policy decisions. He doesn’t want to sit in meetings, trying to form policy juggling people’s egos. That’s where the CSO comes in.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*fNn6CLvkByfEA_z7ZJ4WvQ.jpeg" alt=""/></p>

<p>Let’s take a concrete use case. As I’m sure you’ve heard, the City of San Diego is getting a deployment of Smart Street streetlights. It’s the first and largest Internet of Things deployment in a city in North America.
These streetlight sensors have a lot of different capabilities (blog post coming soon). They have the potential to improve parking management. City staff can use them to monitor pedestrian and car traffic. We can use them to track our environment.</p>

<p>Without reading the API documentation it’s hard to understand what’s possible. It takes someone with a background in software to do that. Then, he would have to translate it to humanese. The CDO happens to enjoy reading that kind of stuff. The CSO definitely does not.</p>

<p>The CDO has a full understanding of the sensors’ capabilities. The CSO has in-depth knowledge of what we need to monitor and change to meet climate goals. Together, we can decide how to utilize this technology to meet her business objectives.</p>

<p>These sensors also need a coordinated policy effort to manage the data. The CSO happens to be good at that as policymaking is part of her job. She can help me push through the policy work that I need.</p>

<p>Our unique skills and expertise complement our shared drive for change and our similar desire to get sh*t done.</p>

<p>The data movement in cities has followed an upward curve similar to the sustainability movement. They have also followed a similar time frame. This is not by accident — both of these movements are new ways of thinking and new policy directions. This is not simple correlation. These roles are very much intertwined and are supporting and enabling each other’s day to day work.</p>
]]></content:encoded>
      <guid>https://bymaksim.com/chief-data-and-sustainability-officers-form-an-effective-team</guid>
      <pubDate>Wed, 28 Feb 2018 13:03:11 +0000</pubDate>
    </item>
    <item>
      <title>Why data automation matters.</title>
      <link>https://bymaksim.com/why-data-automation-matters?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[This article first appeared on data.sandiego.gov&#xA;&#xA;Be honest. When you read the words data automation you get a sudden rush of melatonin to your brain, your eyelids get heavy, and you get an uncontrollable urge to fall asleep. Don’t be ashamed; you are reacting to these words in a similar manner to 99.5% of people on the planet. Bear with me though for just a few paragraphs while I try to explain why it matters, and how we do it here at the City of San Diego.&#xA;&#xA;!--more--&#xA;&#xA;Before we get started, let me promise to not use any of the following:&#xA;&#xA;Mathematics (or any subset thereof)&#xA;&#xA;Buzzwords&#xA;&#xA;Tech Jargon&#xA;&#xA;The Quick Win&#xA;&#xA;One of the best case studies to make obvious the benefits of data automation is StreetsSD. We started working on this project with the City’s Transportation and Stormwater Department (TSW). They asked us to build a map of paving projects. We got a spreadsheet of data from the department, loaded it into our mapping tool and wrote some code. Boom, we had a map.&#xA;&#xA;The Ugly Loss&#xA;&#xA;Most web maps, charts, and data visualizations you see in newspapers, and even those created by a lot of cities, usually stop there. That’s actually OK because a lot of these things don’t need to be updated automatically. However, the goal of StreetsSD is to provide an up-to-date status of work to city employees and residents. Keeping the map updated is important.&#xA;&#xA;The typical City approach in these situations is to have someone “run a report” and upload it somewhere to update the map. Most of the time, “running a report” actually means this:&#xA;&#xA;Extract data from a database, using a query (that could change, resulting in different data).&#xA;&#xA;Complete 1–40 manual steps in Excel to clean the data. These steps are usually not documented anywhere and could easily vary for each run of the report.&#xA;&#xA;Send it or upload it somewhere.&#xA;&#xA;This approach inherently has several problems:&#xA;&#xA;Data is inconsistent, meaning source of truth is inconsistent&#xA;&#xA;Bad data potentially causes bugs in the application&#xA;&#xA;City employees waste a HUGE amount of time&#xA;&#xA;People might forget to actually run the report&#xA;&#xA;We can’t expect someone to run a report every day (or with any kind of frequency that a continously update data source needs).&#xA;&#xA;Needless to say, this sucks. And even though we’re creating a great operational tool, we’re costing people time and risking the release of wrong information.&#xA;&#xA;At this point, being the wise reader you are, you might be thinking that the outcome would definitely be worth the cost in the case of just one map. To which I’ll respond with:&#xA;&#xA;This is not a problem with just one map. It would occur with any dataset we provide as open data on the portal, anything we build on top of those datasets, and any internal or external reports we regularly generate. Therefore, this is a problem that needs to be solved at a systemic level.&#xA;&#xA;Redemption&#xA;&#xA;Our philosophy has always been to let machines do what they do best — updating data, re-running things, keeping track of things — and humans do what they do best — making those fuzzy decisions that only our brains can.&#xA;&#xA;We needed a flexible and extensible solution that could scale across the organization as a whole.&#xA;&#xA;We turned to an open source project called Airflow. It has become the tool of choice for Spotify, IFTTT, Lyft, AirBnB and many others for solving this exact set of problems. Airflow has several advantages:&#xA;&#xA;Built in Python — so it’s extensible with any of thousands of Python open-source packages&#xA;&#xA;Modular — new connections to different data sources are easy to write&#xA;&#xA;Open source with a strong community — so there’s plenty of support and continuous updates&#xA;&#xA;Scalable — as we need to do more and more, Airflow easily scales&#xA;&#xA;We called our Airflow deployment Poseidon because codenames are cool, San Diego is on the Ocean, and Poseidon rules the sea. Plus, we get to use images like these:&#xA;&#xA;The basic idea of Poseidon is this:&#xA;&#xA;Get data from a source (database, spreadsheet, map, website)&#xA;&#xA;Do some stuff to it (geocode it, aggregate it, clean it)&#xA;&#xA;Upload it (put it on the cloud that is backing our portal and a variety of other applications)&#xA;&#xA;Run it on a schedule (once every 5 minutes, OR once every day, OR on every odd day of the month at 3:02 PM)&#xA;&#xA;Simple, right?&#xA;&#xA;We’re now doing this for all of our datasets (way harder to implement than explain). This means we’re also doing it for everything built on top of our data, such as StreetsSD, the portal, and various other visualizations.&#xA;&#xA;The Open Sea&#xA;&#xA;Now you’re probably thinking: whoop-de-doo, I’m a resident and I don’t know data, so I really don’t care that you now have automated data.&#xA;&#xA;We know. That’s why we pushed it further.&#xA;&#xA;Look back at the diagram above. Those are just dependent pieces that run in a defined cycle. What if we flipped out some of those pieces and got this?&#xA;&#xA;Automated data is now starting to get interesting. Because of Poseidon’s flexibility, and the fact that it’s just a bunch of coordinated tasks, we can send automatic notifications, alerts based on thresholds, and all kinds of other cool stuff that would never have been possible without automation.&#xA;&#xA;But if I told you what we were planning next, I’d ruin the surprise. You’ll just have to wait and see.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>This article first appeared on <a href="https://data.sandiego.gov" rel="nofollow">data.sandiego.gov</a></p>

<p>Be honest. When you read the words data automation you get a sudden rush of melatonin to your brain, your eyelids get heavy, and you get an uncontrollable urge to fall asleep. Don’t be ashamed; you are reacting to these words in a similar manner to 99.5% of people on the planet. Bear with me though for just a few paragraphs while I try to explain why it matters, and how we do it here at the City of San Diego.</p>



<p>Before we get started, let me promise to not use any of the following:</p>
<ul><li><p>Mathematics (or any subset thereof)</p></li>

<li><p>Buzzwords</p></li>

<li><p>Tech Jargon</p></li></ul>

<h2 id="the-quick-win">The Quick Win</h2>

<p>One of the best case studies to make obvious the benefits of data automation is <a href="http://streets.sandiego.gov/" rel="nofollow">**StreetsSD</a><strong>. We started working on this project with the City’s Transportation and Stormwater Department (TSW). They asked us to build a map of paving projects. We got a spreadsheet of data from the department, loaded it into [</strong>our mapping tool](<a href="http://carto.com/)**" rel="nofollow">http://carto.com/)**</a> and wrote some code. Boom, we had a map.</p>

<p><img src="https://cdn-images-1.medium.com/max/2536/1*lAOFYk_zQEWF9AO2g6666A.jpeg" alt=""/></p>

<h2 id="the-ugly-loss">The Ugly Loss</h2>

<p>Most web maps, charts, and data visualizations you see in newspapers, and even those created by a lot of cities, usually stop there. That’s actually OK because a lot of these things don’t need to be updated automatically. However, the goal of StreetsSD is to provide an up-to-date status of work to city employees and residents. Keeping the map updated is important.</p>

<p>The typical City approach in these situations is to have someone “run a report” and upload it somewhere to update the map. Most of the time, “running a report” actually means this:</p>
<ul><li><p>Extract data from a database, using a query (that could change, resulting in different data).</p></li>

<li><p>Complete 1–40 manual steps in Excel to clean the data. These steps are usually not documented anywhere and could easily vary for each run of the report.</p></li>

<li><p>Send it or upload it somewhere.</p></li></ul>

<p>This approach inherently has several problems:</p>
<ul><li><p>Data is inconsistent, meaning source of truth is inconsistent</p></li>

<li><p>Bad data potentially causes bugs in the application</p></li>

<li><p>City employees waste a HUGE amount of time</p></li>

<li><p>People might forget to actually run the report</p></li>

<li><p>We can’t expect someone to run a report every day (or with any kind of frequency that a continously update data source needs).</p></li></ul>

<p>Needless to say, this sucks. And even though we’re creating a great operational tool, we’re costing people time and risking the release of wrong information.</p>

<p>At this point, being the wise reader you are, you might be thinking that the outcome would definitely be worth the cost in the case of just one map. To which I’ll respond with:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*KxL-9aJtRiDfL8XfnXzlwQ.jpeg" alt=""/></p>

<p>This is not a problem with just one map. It would occur with any dataset we provide as open data on the portal, anything we build on top of those datasets, and any internal or external reports we regularly generate. Therefore, this is a problem that needs to be solved at a systemic level.</p>

<h2 id="redemption">Redemption</h2>

<p>Our philosophy has always been to let machines do what they do best — updating data, re-running things, keeping track of things — and humans do what they do best — making those fuzzy decisions that only our brains can.</p>

<p>We needed a flexible and extensible solution that could scale across the organization as a whole.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*_HD4Ge-SDwg5lOMLRHNSFw.jpeg" alt=""/></p>

<p>We turned to an open source project called <a href="https://github.com/apache/incubator-airflow/" rel="nofollow">**Airflow</a>**. It has become the tool of choice for Spotify, IFTTT, Lyft, AirBnB and many others for solving this exact set of problems. Airflow has several advantages:</p>
<ul><li><p>Built in Python — so it’s extensible with any of thousands of Python open-source packages</p></li>

<li><p>Modular — new connections to different data sources are easy to write</p></li>

<li><p>Open source with a strong community — so there’s plenty of support and continuous updates</p></li>

<li><p>Scalable — as we need to do more and more, Airflow easily scales</p></li></ul>

<p>We called our Airflow deployment Poseidon because codenames are cool, San Diego is on the Ocean, and Poseidon rules the sea. Plus, we get to use images like these:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*zYn0nriHJR4xHku1uz6OHw.jpeg" alt=""/></p>

<p>The basic idea of Poseidon is this:</p>

<p><img src="https://cdn-images-1.medium.com/max/2138/1*tdKnXa9ksRnu2MwJBzuzfQ.jpeg" alt=""/></p>
<ul><li><p>Get data from a source (database, spreadsheet, map, website)</p></li>

<li><p>Do some stuff to it (geocode it, aggregate it, clean it)</p></li>

<li><p>Upload it (put it on the cloud that is backing our portal and a variety of other applications)</p></li>

<li><p>Run it on a schedule (once every 5 minutes, OR once every day, OR on every odd day of the month at 3:02 PM)</p></li></ul>

<p>Simple, right?</p>

<p>We’re now doing this for all of our datasets (way harder to implement than explain). This means we’re also doing it for everything built on top of our data, such as StreetsSD, the portal, and various other visualizations.</p>

<h2 id="the-open-sea">The Open Sea</h2>

<p>Now you’re probably thinking: whoop-de-doo, I’m a resident and I don’t know data, so I really don’t care that you now have automated data.</p>

<p>We know. That’s why we pushed it further.</p>

<p>Look back at the diagram above. Those are just dependent pieces that run in a defined cycle. What if we flipped out some of those pieces and got this?</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*7dZrpyA6xNm5X-k9rxuGlQ.jpeg" alt=""/></p>

<p>Automated data is now starting to get interesting. Because of Poseidon’s flexibility, and the fact that it’s just a bunch of coordinated tasks, we can send automatic notifications, alerts based on thresholds, and all kinds of other cool stuff that would never have been possible without automation.</p>

<p>But if I told you what we were planning next, I’d ruin the surprise. You’ll just have to wait and see.</p>
]]></content:encoded>
      <guid>https://bymaksim.com/why-data-automation-matters</guid>
      <pubDate>Thu, 27 Apr 2017 12:05:34 +0000</pubDate>
    </item>
    <item>
      <title>A Faster San Diego Data Portal</title>
      <link>https://bymaksim.com/a-faster-san-diego-data-portal?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[&#xA;Hey San Diego! Your open data portal just got a LOT faster!&#xA;&#xA;One of the reasons we wanted to run our own data portal is the flexibility we have to change it and add functionality.&#xA;&#xA;Today, we’re putting the pedal to the metal on those desires. We initially launched the portal based on JKAN, but with modified schemas, layouts, and branding. Because of how fast we moved, we put off thinking about speed and performance.&#xA;&#xA;!--more--&#xA;&#xA;Since the dust settled a bit, we had a chance to do that.&#xA;&#xA;Lighthouse is a tool developed by Google to test web pages for performance, accessibility, and more. The first time we ran it against our portal, here’s what we got:&#xA;&#xA;See full report&#xA;&#xA;30 out of 100, with a lot of marks against the portal due to loading speed. Time to first paint is a metric that measures when the primary content of a page is visible. That number means it took 5.3 seconds after you navigated to the portal for anything to show up.&#xA;&#xA;We did badly on plenty of other metrics, including SSL encryption (HTTPS) and offline browsing.&#xA;&#xA;So why does this matter (besides our obsessive perfectionism)? KissMetrics has a great infographic about this:&#xA;&#xA;In short, we were losing an estimated 25 percent of portal users because of the page loading speed. Not good.&#xA;&#xA;My friend and former Code for America co-fellow David Leonard (who also got me into Polymer components that the portal heavily uses ) helped me analyze where the bottlenecks were in our page loading time and gave me some tips about how to speed up the portal.&#xA;&#xA;We implemented https and service workers (for security and offline browsing and caching) and optimized how components load. The portal is now encrypted, and you can browse it offline.&#xA;&#xA;I’m not going to go into too much detail, but you can see the pull request here. We tested again and got a score of 96:&#xA;&#xA;See full report&#xA;&#xA;Let me use a few buzzwords to describe what happened here, for those fond of them:&#xA;&#xA;We used an agile process to iterate on the technology underlying the performance, availability, and scalability of our portal architecture to decrease page loading speed, increase cybersecurity, and improve an already great product with enhanced user experience.&#xA;&#xA;Or in English:&#xA;&#xA;We did a thing that most organizations do: we improved something we launched, because no one ever gets everything right the first time. We’ll do it again, and again, and again. This is just the first of many changes and enhancements we will be making.*&#xA;&#xA;Enjoy your [faster] portal, San Diego!&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>Hey San Diego! Your open data portal just got a LOT faster!</p>

<p>One of the reasons we wanted to <a href="https://data.sandiego.gov/stories/portal-refresh" rel="nofollow">run our own data portal</a> is the flexibility we have to change it and add functionality.</p>

<p>Today, we’re putting the pedal to the metal on those desires. We initially launched the portal based on JKAN, but with modified schemas, layouts, and branding. Because of how fast we moved, we put off thinking about speed and performance.</p>



<p>Since the dust settled a bit, we had a chance to do that.</p>

<p><a href="https://developers.google.com/web/tools/lighthouse/" rel="nofollow">Lighthouse</a> is a tool developed by Google to test web pages for performance, accessibility, and more. The first time we ran it against our portal, here’s what we got:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*I8sLRtUaR3egDzhtJixVNw.jpeg" alt=""/></p>

<p><a href="http://theia.datasd.org.s3.amazonaws.com/data.sandiego.gov_2017-03-02_19-26-42.html" rel="nofollow">See full report</a></p>

<p>30 out of 100, with a lot of marks against the portal due to loading speed. Time to first paint is a metric that measures when the primary content of a page is visible. That number means it took 5.3 seconds after you navigated to the portal for anything to show up.</p>

<p>We did badly on plenty of other metrics, including SSL encryption (HTTPS) and offline browsing.</p>

<p>So why does this matter (besides our obsessive perfectionism)? KissMetrics has a <a href="https://blog.kissmetrics.com/loading-time/?wide=1" rel="nofollow">great infographic about this</a>:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*A1Cb15-sFRUPC3LBB-cC2Q.jpeg" alt=""/></p>

<p>In short, we were losing an estimated 25 percent of portal users because of the page loading speed. Not good.</p>

<p>My friend and former Code for America co-fellow <a href="http://twitter.com/davidleonardii" rel="nofollow">David Leonard</a> (who also got me into Polymer components that the <a href="https://data.sandiego.gov/stories/portal-refresh" rel="nofollow">portal heavily uses</a> ) helped me analyze where the bottlenecks were in our page loading time and gave me some tips about how to speed up the portal.</p>

<p>We implemented https and service workers (for security and offline browsing and caching) and optimized how components load. The portal is now encrypted, and you can browse it offline.</p>

<p>I’m not going to go into too much detail, but you can <a href="https://github.com/cityofsandiego/seaboard/pull/140" rel="nofollow">see the pull request here</a>. We tested again and got a score of 96:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/1*nhE802QN8XbPb0umssNTEA.jpeg" alt=""/></p>

<p><a href="http://theia.datasd.org.s3.amazonaws.com/data.sandiego.gov_2017-04-12_18-17-58.html" rel="nofollow">See full report</a></p>

<p>Let me use a few buzzwords to describe what happened here, for those fond of them:</p>

<p><strong><em>We used an agile process to iterate on the technology underlying the performance, availability, and scalability of our portal architecture to decrease page loading speed, increase cybersecurity, and improve an already great product with enhanced user experience.</em></strong></p>

<p>Or in English:</p>

<p><strong><em>We did a thing that most organizations do: we improved something we launched, because no one ever gets everything right the first time. We’ll do it again, and again, and again. This is just the first of many changes and enhancements we will be making.</em></strong></p>

<p>Enjoy your [faster] portal, San Diego!</p>
]]></content:encoded>
      <guid>https://bymaksim.com/a-faster-san-diego-data-portal</guid>
      <pubDate>Thu, 20 Apr 2017 00:30:59 +0000</pubDate>
    </item>
    <item>
      <title>Savvy Savannah</title>
      <link>https://bymaksim.com/savvy-savannah?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[The hackathon was focused around conservation, and preventing poaching, with some very explicit problem statements. The issue we hoped to address was to minimize poaching of large animals such as Rhinos and Elephants in wildlife preserves around the world.&#xA;&#xA;Please see the DevPost Project Page for more info about the project, the hackathon and the team.&#xA;&#xA;!--more--&#xA;&#xA;Inspiration&#xA;&#xA;Endangered species in conservation areas are threatened to the point of extinction by poachers. Local authorities are resource constrained and are not able to monitor an entire conservation area for poaching activity — particularly the interior of a preserve. Humans in general would not be very good at this either because of the high-friction way reporting would have to be done. Our solution was deploy low-cost connected sensors throughout the interior of a conservation area to monitor the area for suspicious activity.&#xA;&#xA;What it Does&#xA;&#xA;Savvy Savanna allows low cost IoT — the sensor and the button — to be deployed in conservation areas and surrounding villages. The system allows people to press a button to report suspicious poaching activity and solar-powered field sensors which automatically report potential poacher problems in the savanna.&#xA;&#xA;Sensors in the savanna detect sounds, such as helicopter noise, that trigger an alert to be sent to the cloud and subsequently to the appropriate authorities.&#xA;&#xA;The system alerts authorities of poacher activity detected by the system. The authorities receive a text message alert and a location of where to respond.&#xA;&#xA;Challenges&#xA;&#xA;Cellular signal reception and WiFi network coverage is limited in the savanna. The sensors need a way to transmit information to the cloud. To solve this problem, the sensors in the field are part of a mesh network. This is technology currently utilized in “smart cities” with interconnected devices. Field sensors can be networked together with only certain sensors directly connected to the internet network.&#xA;&#xA;Community members may be targeted for attack by organized crime groups if they are seen as reporting poaching activities. To solve this problem, “the button” component of the system is a way that community members can anonymously report potential poaching activity in their neighborhood by simply pushing a button.&#xA;&#xA;Savvy Overview Prezi&#xA;&#xA;Cost Estimation&#xA;&#xA;Our rough estimates are that with the current architecture, deploying one of these one per 0.2 square miles would cost around $20,000. However, as we make changes to the hardware and implement mesh networking, this cost could significantly go down.&#xA;&#xA;How we Built The Sound Sensor&#xA;&#xA;Hardware&#xA;&#xA;The hardware consists of a Particle Electron Board.&#xA; These are Arduino based boards that have a connection to the internet through a 2G network available in over 100 countries in the world.&#xA;&#xA;Particle ElectronParticle Electron&#xA;&#xA;They are backed up by Particle’s cloud infrastructure. This allows us to trigger webhooks directly using the firmware on the board. We can also use this architecture to expose different functions on the board as POST requests, allowing us to build an API for our IoT device.&#xA;&#xA;We started by wiring up the Particle board with:  Analog microphone  LED for monitoring the status of the run loop (debugging)  LED for monitoring the input volume of the sound coming into the mic&#xA;&#xA;There is an additional LED on the electron board itself. We used that to communicate that the sound threshold has been reached an a notification is being triggered.&#xA;&#xA;Particle Wire UpParticle Wire Up&#xA;&#xA;What Happens&#xA;&#xA;The sensor sits there and listens. Arduino devices operate on a loop, so it samples the sound volume every 2 seconds. Since it’s an analog input, the volume will be anywhere from 0 to 4096. When the volume spikes over 2000, a command is sent to Particle Cloud. This triggers a POST webhook with the data sent from Particle to AWS API Gateway which then relays the input to AWS Lambda. The Lambda callback function then triggers the Twilio API to deploy text messages.&#xA;&#xA;Lastly, we use the MapBox static maps api (you can use API to generate a static map image with the location and send that through as Twilio MMS. Since we don’t have a GPS sensor on the button or the Electron just yet, we used StaticMapGen to prototype.&#xA;&#xA;How we Built The AWS IoT Button&#xA;&#xA;Hardware&#xA;&#xA;While a sound sensor is nice to have for continous monitoring, our idea was also to have a way for community members to anonymously report poaching, avoiding retaliation from organized crime groups. We used an AWS button as a stand-in. However, since it’s only wi-fi enabled and carries no gps chip, ideally we would build a custom one.&#xA;&#xA;What happens&#xA;&#xA;Because of the modular architecture we already had in place it was fairly simple to add the AWS IoT trigger as an additional trigger for the Lambda function.&#xA;&#xA;How it all fits together:&#xA;&#xA;Overall ArchitectureOverall Architecture&#xA;&#xA;Code&#xA;&#xA;Grab the Arduino code for the firmware&#xA;&#xA;Grab the Lambda Function JS&#xA;&#xA;Next Steps&#xA;&#xA;We will probably continue hacking on this project, as this is an important challenge and we think we may have a way to solve it. We still have quite a bit of ways to go, specifically:&#xA;&#xA;GPS Enable the sensors&#xA;&#xA;Solar Enable the sensors&#xA;&#xA;Provide protective covering for the sensors&#xA;&#xA;Provide cloud infrastructure for data aggregation and monitoring&#xA;&#xA;Train a model to recognize sound (elephant trumpet vs helicopter noise) and deploy it to the devices on the edge. This may require switching to integrating something like the Intel Edison chip.&#xA;&#xA;Switch devices to using a low-power WAN configuration, so they can operate in a mesh with only a few points of direct network connectivity to the web.&#xA;&#xA;Implement something like Carloop.io in ranger vehicles to prevent false positives caused by ranger vehicle noise.&#xA;&#xA;Keep track of what we’re up to and jump in over in the Github Issue Queue&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>The hackathon was focused around conservation, and preventing poaching, with some very explicit problem statements. The issue we hoped to address was to minimize poaching of large animals such as Rhinos and Elephants in wildlife preserves around the world.</p>

<p>Please see the <a href="https://devpost.com/software/savvy-savanna" rel="nofollow">DevPost Project Page</a> for more info about the project, the hackathon and the team.</p>



<h3 id="inspiration">Inspiration</h3>

<p>Endangered species in conservation areas are threatened to the point of extinction by poachers. Local authorities are resource constrained and are not able to monitor an entire conservation area for poaching activity — particularly the interior of a preserve. Humans in general would not be very good at this either because of the high-friction way reporting would have to be done. Our solution was deploy low-cost connected sensors throughout the interior of a conservation area to monitor the area for suspicious activity.</p>

<h3 id="what-it-does">What it Does</h3>

<p>Savvy Savanna allows low cost IoT — the sensor and the button — to be deployed in conservation areas and surrounding villages. The system allows people to press a button to report suspicious poaching activity and solar-powered field sensors which automatically report potential poacher problems in the savanna.</p>

<p>Sensors in the savanna detect sounds, such as helicopter noise, that trigger an alert to be sent to the cloud and subsequently to the appropriate authorities.</p>

<p>The system alerts authorities of poacher activity detected by the system. The authorities receive a text message alert and a location of where to respond.</p>

<h3 id="challenges">Challenges</h3>

<p>Cellular signal reception and WiFi network coverage is limited in the savanna. The sensors need a way to transmit information to the cloud. To solve this problem, the sensors in the field are part of a mesh network. This is technology currently utilized in “smart cities” with interconnected devices. Field sensors can be networked together with only certain sensors directly connected to the internet network.</p>

<p>Community members may be targeted for attack by organized crime groups if they are seen as reporting poaching activities. To solve this problem, “the button” component of the system is a way that community members can anonymously report potential poaching activity in their neighborhood by simply pushing a button.</p>

<p><a href="https://prezi.com/embed/t9nosx2i5ey0/?bgcolor=ffffff&amp;amp;lock_to_path=0&amp;amp;autoplay=0&amp;amp;autohide_ctrls=0&amp;amp;landing_data=bHVZZmNaNDBIWnNjdEVENDRhZDFNZGNIUE43MHdLNWpsdFJLb2ZHanI5dWRlclpwaUwwaG1sd0N1VHNzMW9wVnFnPT0&amp;amp;landing_sign=aR6rnRpwpxMhRosjUunK9Iud-CZ-sUipNWlRRk2zGk4" rel="nofollow">Savvy Overview Prezi</a></p>

<h3 id="cost-estimation">Cost Estimation</h3>

<p>Our rough estimates are that with the current architecture, deploying one of these one per 0.2 square miles would cost around $20,000. However, as we make changes to the hardware and implement mesh networking, this cost could significantly go down.</p>

<h2 id="how-we-built-the-sound-sensor">How we Built The Sound Sensor</h2>

<h3 id="hardware">Hardware</h3>

<p>The hardware consists of a Particle Electron Board.
 These are Arduino based boards that have a connection to the internet through a 2G network available in over 100 countries in the world.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*3q8Pw2guBbXRBUDl." alt="Particle Electron"/><em>Particle Electron</em></p>

<p>They are backed up by Particle’s cloud infrastructure. This allows us to trigger webhooks directly using the firmware on the board. We can also use this architecture to expose different functions on the board as POST requests, allowing us to build an API for our IoT device.</p>

<p>We started by wiring up the Particle board with: * Analog microphone * LED for monitoring the status of the run loop (debugging) * LED for monitoring the input volume of the sound coming into the mic</p>

<p>There is an additional LED on the electron board itself. We used that to communicate that the sound threshold has been reached an a notification is being triggered.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*w8B1oVqtHSyRHySb." alt="Particle Wire Up"/><em>Particle Wire Up</em></p>

<h3 id="what-happens">What Happens</h3>

<p>The sensor sits there and listens. Arduino devices operate on a loop, so it samples the sound volume every 2 seconds. Since it’s an analog input, the volume will be anywhere from 0 to 4096. When the volume spikes over 2000, a command is sent to Particle Cloud. This triggers a POST webhook with the data sent from Particle to AWS API Gateway which then relays the input to AWS Lambda. The Lambda callback function then triggers the Twilio API to deploy text messages.</p>

<p>Lastly, we use the MapBox static maps api (you can use API to generate a static map image with the location and send that through as Twilio MMS. Since we don’t have a GPS sensor on the button or the Electron just yet, we used <a href="http://staticmapmaker.com/" rel="nofollow">StaticMapGen</a> to prototype.</p>

<h2 id="how-we-built-the-aws-iot-button">How we Built The AWS IoT Button</h2>

<h3 id="hardware-1">Hardware</h3>

<p>While a sound sensor is nice to have for continous monitoring, our idea was also to have a way for community members to anonymously report poaching, avoiding retaliation from organized crime groups. We used an AWS button as a stand-in. However, since it’s only wi-fi enabled and carries no gps chip, ideally we would build a custom one.</p>

<h3 id="what-happens-1">What happens</h3>

<p>Because of the modular architecture we already had in place it was fairly simple to add the AWS IoT trigger as an additional trigger for the Lambda function.</p>

<h3 id="how-it-all-fits-together">How it all fits together:</h3>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*L5R6uCurfaKSmA1L." alt="Overall Architecture"/><em>Overall Architecture</em></p>

<h3 id="code">Code</h3>

<p>Grab the <a href="https://github.com/MrMaksimize/whistler" rel="nofollow">Arduino code for the firmware</a></p>

<p>Grab the <a href="https://github.com/MrMaksimize/whistler_backend" rel="nofollow">Lambda Function JS</a></p>

<h3 id="next-steps">Next Steps</h3>

<p>We will probably continue hacking on this project, as this is an important challenge and we think we may have a way to solve it. We still have quite a bit of ways to go, specifically:</p>
<ul><li><p>GPS Enable the sensors</p></li>

<li><p>Solar Enable the sensors</p></li>

<li><p>Provide protective covering for the sensors</p></li>

<li><p>Provide cloud infrastructure for data aggregation and monitoring</p></li>

<li><p>Train a model to recognize sound (elephant trumpet vs helicopter noise) and deploy it to the devices on the edge. This may require switching to integrating something like the Intel Edison chip.</p></li>

<li><p>Switch devices to using a low-power WAN configuration, so they can operate in a mesh with only a few points of direct network connectivity to the web.</p></li>

<li><p>Implement something like <a href="http://carloop.io" rel="nofollow">Carloop.io</a> in ranger vehicles to prevent false positives caused by ranger vehicle noise.</p></li></ul>

<p>Keep track of what we’re up to and jump in over in the <a href="https://github.com/MrMaksimize/whistler/issues" rel="nofollow">Github Issue Queue</a></p>
]]></content:encoded>
      <guid>https://bymaksim.com/savvy-savannah</guid>
      <pubDate>Fri, 11 Nov 2016 01:28:44 +0000</pubDate>
    </item>
    <item>
      <title>StreetsSD Overview</title>
      <link>https://bymaksim.com/streetssd-overview?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[StreetsSD was an interesting project us from an organizational and technical perspective. Let’s peek behind the scenes to see how this all came together.&#xA;&#xA;!--more--&#xA;&#xA;Overall Architecture&#xA;&#xA;Jekyll&#xA;&#xA;Overall, StreetsSD is a Jekyll site. This allows us to avoid maintenance costs, performance issues, and complexity associated with running database-driven sites. We host it using Github Pages for the wonderful price of $0.&#xA;&#xA;Jekyll OutlineJekyll Outline&#xA;&#xA;Everything you see outlined in pink above is Jekyll and custom JS. We used the San Diego Style Guide for colors and layouts. The charts on the bottom, the layers, and the text in the explainer boxes are controlled by the views.yml in the repo.&#xA;&#xA;All the dynamic components you see from the layers (FY14, FY15, etc) to the charts and totals are populated through custom SQL queries against data in Carto. They are constructed using the squel library in the sqlBuilder file.&#xA;&#xA;As an example, when you click on the “Work in FY–2016” layer, an event is sent that constructs and sends the query to Carto to display results on a map. Based on the views.yml file, we know that layer has WorkType and WorkByMonth charts, so one more query is constructed and sent to Carto. After additional JS data processing, we are able to display the charts.&#xA;&#xA;Carto&#xA;&#xA;The interactive map itself is all Carto. Most of the visual components (e.g. line styles, legends, hovers) are stored in Carto as well. We don’t love that, because those should be stored in code, but we have an issue on the backlog to address it.&#xA;&#xA;Carto editorCarto editor&#xA;&#xA;The map in Carto has two “layers” — one for OCI and one for Streetwork, both based off sdstreetsquery table. Carto’s model for associating layers to datasets ties them together permanently. By having a virtual dataset that we can impose queries on, we are eliminating the tie between a dataset and a map, giving us more flexibility to flip out the underlying datasets if we need to change schemas.&#xA;&#xA;Then why two layers from one table? Because OCI layers and streetwork layers have slightly different fields and labels. Carto’s interactive UI is tied to a layer, and we had to separate the two. Then, we impose the queries on the layer accordingly. We can probably condense this to one layer using more advanced conditional logic, but we’ll stay away from that for now.&#xA;&#xA;Deployment&#xA;&#xA;We use the master branch as the bleeding edge. That gets deployed to a staging site continuously by CircleCI. The production branch gets deployed to the gh-pages branch, which is what feeds StreetsSD. We get notifications about each step.&#xA;&#xA;Datasets&#xA;&#xA;cgstreetscombined&#xA;&#xA;We created this dataset by merging the City’s official street file geographies with essential data from the City’s Pavement Management System. Every other set of data is joined against cgstreetscombined by segmentid to derive the geometries and street names.&#xA;&#xA;sdpavingdatasd&#xA;&#xA;This is the full work layer, encompassing work done from the beginning of FY14. We’ll talk about the specifics in the data delivery section. This has no geospatial data. As layers get selected, SQL is generated, applied to this layer, and joined against cgstreetscombined.&#xA;&#xA;oci2011datasd and oci2015datasd&#xA;&#xA;These were created using historical data provided in the Pavement Management System and the new OCI numbers we received from latest run of the most recent vendor. These also have no geospatial data. As layers get selected, SQL is generated, applied to this layer, and joined against cgstreetscombined.&#xA;&#xA;Data Discovery&#xA;&#xA;We partnered with the Transportation and Storm Water Department to build the map. However, we quickly ran into some challenges with understanding the data, and the calculations performed for reporting.&#xA;&#xA;The PVM (let’s use that as the acronym for the Pavement Management System) is an application that is built on top of SQL server using queries and stored procedures. Every month, the streets team has to update IMCAT. IMCAT is a system the City uses to coordinate street work across departments. This way we can avoid the same segment being dug up multiple times in a short time span. The streets team pull a report from the PVM generated by a query and a stored procedure). Then, it’s manually cleaned in Excel (we counted 32 steps), joined with a streets shapefile, and uploaded it to IMCAT.&#xA;&#xA;This is where we started. The data that gets uploaded to IMCAT has to be in a particular format with certain columns, so we couldn’t directly trace that dataset to the query running against PVM. And we couldn’t just guess the query and match the result either — PVM has 319 tables.&#xA;&#xA;A few months ago, the City invested in a tool called Alation exactly for this type of use case. We wanted to be able to maintain data catalogs and keep documentation on the data that we have, allowing data users to be better informed. This database is the first one we connected with Alation.&#xA;&#xA;We built a robust query to continuously pull data. To do this, we used Alation to trace the query the PVM application was running. However, we had no visibility into the stored procedure. To bypass this, we combined information from using the query, the streets team and Alation’s query logs.&#xA;&#xA;Query LogQuery Log&#xA;&#xA;TablesTables&#xA;&#xA;Table Internal View + Our NotesTable Internal View + Our Notes&#xA;&#xA;Built Query for Paving DataBuilt Query for Paving Data&#xA;&#xA;Built Query for Street InfoBuilt Query for Street Info&#xA;&#xA;Going through this process gave us an enormous amount of knowledge of street paving information, which we have documented in Alation.&#xA;&#xA;However, we weren’t done yet. We still needed to build the ETL for Street Map Data, IMCAT data (because friends don’t let friends clean data over and over manually), and the CGSTREETSCOMBINED file.&#xA;&#xA;Data Delivery&#xA;&#xA;We have two main queries that pull data without doing much cleaning. We wanted to stay away from cleaning in SQL as a best practice.&#xA;&#xA;streetsbase gets pulled by FME, combined with the geospatial base file, and outputs as a shapefile to S3. Carto syncs it down from there:&#xA;&#xA;Streets Base ETLStreets Base ETL&#xA;&#xA;R pulls Pavement Ex in using Alation’s API and the Alation R Package, completes the 32 manual steps the department normally has to do (with some conditional forks for IMCAT), and then outputs the sdpavingdatasd.csv file into S3. We also make sure to standardize the outgoing data according to our Technical Guidelines. At the same time, we output the IMCAT file to be ingested for the conflict mitigation map.&#xA;&#xA;Verifying Calculations&#xA;&#xA;So, how in the world do we know that our constructed SQL on the client side is doing the right thing and making the right calculations? We use R to verify that the calculations are correct by pulling the data straight from Alation, and duplicating the calculations that the SQL does on Streets Map.&#xA;&#xA;What’s next?&#xA;&#xA;This is an alpha project, and we plan to continue working on it. We will work in sprints. As people submit feature or bug requests, we will prioritize them according to what we think we can do during the time allotted. Critical bugs always get fixed first.&#xA;&#xA;Of course, we are more than open to Pull Requests. If there is a feature you really want and we’re just unable to get to it, feel free to discuss it with us in the issue queues, fork the code, and make a Pull Request.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>StreetsSD was an interesting project us from an organizational and technical perspective. Let’s peek behind the scenes to see how this all came together.</p>



<h2 id="overall-architecture">Overall Architecture</h2>

<h3 id="jekyll">Jekyll</h3>

<p>Overall, <a href="http://streets.sandiego.gov" rel="nofollow">StreetsSD</a> is a <a href="https://jekyllrb.com/" rel="nofollow">Jekyll</a> site. This allows us to avoid maintenance costs, performance issues, and complexity associated with running database-driven sites. We host it using Github Pages for the wonderful price of $0.</p>

<p><img src="https://cdn-images-1.medium.com/max/2574/0*8eE46LOTUjPuwk6K." alt="Jekyll Outline"/><em>Jekyll Outline</em></p>

<p>Everything you see outlined in pink above is Jekyll and custom JS. We used the <a href="www.sandiego.gov/communications/design/index.shtml" rel="nofollow">San Diego Style Guide</a> for colors and layouts. The charts on the bottom, the layers, and the text in the explainer boxes are controlled by the <a href="https://github.com/cityofsandiego/streetsSD/blob/master/src/_data/views.yml" rel="nofollow">views.yml</a> in the repo.</p>

<p>All the dynamic components you see from the layers (FY14, FY15, etc) to the charts and totals are populated through custom SQL queries against data in Carto. They are constructed using the <a href="https://hiddentao.com/squel/" rel="nofollow">squel library</a> in the <a href="https://github.com/cityofsandiego/streetsSD/blob/master/src/assets/javascript/sqlBuilder.js" rel="nofollow">sqlBuilder</a> file.</p>

<p>As an example, when you click on the “Work in FY–2016” layer, an event is sent that constructs and sends the query to Carto to display results on a map. Based on the <a href="https://github.com/cityofsandiego/streetsSD/blob/master/src/_data/views.yml" rel="nofollow">views.yml</a> file, we know that layer has WorkType and WorkByMonth charts, so one more query is constructed and sent to Carto. After additional JS data processing, we are able to display the charts.</p>

<h3 id="carto">Carto</h3>

<p>The interactive map itself is all Carto. Most of the visual components (e.g. line styles, legends, hovers) are stored in Carto as well. We don’t love that, because those should be stored in code, but we have an issue on the backlog to address it.</p>

<p><img src="https://cdn-images-1.medium.com/max/2536/0*qsx2XPDhhLY_opiT." alt="Carto editor"/><em>Carto editor</em></p>

<p>The map in Carto has two “layers” — one for OCI and one for Streetwork, both based off sdstreets_query table. Carto’s model for associating layers to datasets ties them together permanently. By having a virtual dataset that we can impose queries on, we are eliminating the tie between a dataset and a map, giving us more flexibility to flip out the underlying datasets if we need to change schemas.</p>

<p>Then why two layers from one table? Because OCI layers and streetwork layers have slightly different fields and labels. Carto’s interactive UI is tied to a layer, and we had to separate the two. Then, we impose the queries on the layer accordingly. We can probably condense this to one layer using more advanced conditional logic, but we’ll stay away from that for now.</p>

<h3 id="deployment">Deployment</h3>

<p>We use the master branch as the bleeding edge. That gets deployed to a staging site continuously by CircleCI. The production branch gets deployed to the gh-pages branch, which is what feeds <a href="http://streets.sandiego.gov" rel="nofollow">StreetsSD</a>. We get notifications about each step.</p>

<h3 id="datasets">Datasets</h3>

<p>cg<em>streets</em>combined</p>

<p>We created this dataset by merging the City’s official street file geographies with essential data from the City’s Pavement Management System. Every other set of data is joined against cg<em>streets</em>combined by segment_id to derive the geometries and street names.</p>

<p>sd<em>paving</em>datasd</p>

<p>This is the full work layer, encompassing work done from the beginning of FY14. We’ll talk about the specifics in the data delivery section. This has no geospatial data. As layers get selected, SQL is generated, applied to this layer, and joined against cg<em>streets</em>combined.</p>

<p>oci<em>2011</em>datasd and oci<em>2015</em>datasd</p>

<p>These were created using historical data provided in the Pavement Management System and the new OCI numbers we received from latest run of the most recent vendor. These also have no geospatial data. As layers get selected, SQL is generated, applied to this layer, and joined against cg<em>streets</em>combined.</p>

<h3 id="data-discovery">Data Discovery</h3>

<p>We partnered with the <a href="https://www.sandiego.gov/tsw" rel="nofollow">Transportation and Storm Water Department</a> to build the map. However, we quickly ran into some challenges with understanding the data, and the calculations performed for reporting.</p>

<p>The PVM (let’s use that as the acronym for the Pavement Management System) is an application that is built on top of SQL server using queries and stored procedures. Every month, the streets team has to update IMCAT. IMCAT is a system the City uses to coordinate street work across departments. This way we can avoid the same segment being dug up multiple times in a short time span. The streets team pull a report from the PVM generated by a query and a stored procedure). Then, it’s manually cleaned in Excel (we counted 32 steps), joined with a streets shapefile, and uploaded it to IMCAT.</p>

<p>This is where we started. The data that gets uploaded to IMCAT has to be in a particular format with certain columns, so we couldn’t directly trace that dataset to the query running against PVM. And we couldn’t just guess the query and match the result either — PVM has 319 tables.</p>

<p>A few months ago, the City invested in a tool called Alation exactly for this type of use case. We wanted to be able to maintain data catalogs and keep documentation on the data that we have, allowing data users to be better informed. This database is the first one we connected with Alation.</p>

<p>We built a robust query to continuously pull data. To do this, we used Alation to trace the query the PVM application was running. However, we had no visibility into the stored procedure. To bypass this, we combined information from using the query, the streets team and Alation’s query logs.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*D4YPkBPgItSpSqx5.png" alt="Query Log"/><em>Query Log</em></p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*_wiQvzgR5QNI3cq3.png" alt="Tables"/><em>Tables</em></p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*lGD_gW8cUglRVtzF.png" alt="Table Internal View + Our Notes"/><em>Table Internal View + Our Notes</em></p>

<p><img src="https://cdn-images-1.medium.com/max/2522/0*GauDEnxUxEd3grBw.png" alt="Built Query for Paving Data"/><em>Built Query for Paving Data</em></p>

<p><img src="https://cdn-images-1.medium.com/max/2534/0*gUlMBn-zyT4psRl7." alt="Built Query for Street Info"/><em>Built Query for Street Info</em></p>

<p>Going through this process gave us an enormous amount of knowledge of street paving information, which we have documented in Alation.</p>

<p>However, we weren’t done yet. We still needed to build the ETL for Street Map Data, IMCAT data (because friends don’t let friends clean data over and over manually), and the CG<em>STREETS</em>COMBINED file.</p>

<h3 id="data-delivery">Data Delivery</h3>

<p>We have two main queries that pull data without doing much cleaning. We wanted to stay away from cleaning in SQL as a best practice.</p>

<p>streets_base gets pulled by FME, combined with the geospatial base file, and outputs as a shapefile to S3. Carto syncs it down from there:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*W8ms2izBY7m3S9I-." alt="Streets Base ETL"/><em>Streets Base ETL</em></p>

<p>R pulls Pavement Ex in using Alation’s API and the <a href="https://github.com/mattwg/alation" rel="nofollow">Alation R Package</a>, completes the 32 manual steps the department normally has to do (with some conditional forks for IMCAT), and then outputs the sd<em>paving</em>datasd.csv file into S3. We also make sure to standardize the outgoing data according to our <a href="https://datasd.gitbooks.io/open-data-implementation-update-2016/content/main/technical-guidelines.html" rel="nofollow">Technical Guidelines</a>. At the same time, we output the IMCAT file to be ingested for the conflict mitigation map.</p>

<h3 id="verifying-calculations">Verifying Calculations</h3>

<p>So, how in the world do we know that our constructed SQL on the client side is doing the right thing and making the right calculations? We use R to verify that the calculations are correct by pulling the data straight from Alation, and duplicating the calculations that the SQL does on Streets Map.</p>

<h2 id="what-s-next">What’s next?</h2>

<p>This is an alpha project, and we plan to continue working on it. We will work in sprints. As people submit feature or bug requests, we will prioritize them according to what we think we can do during the time allotted. Critical bugs always get fixed first.</p>

<p>Of course, we are more than open to Pull Requests. If there is a feature you really want and we’re just unable to get to it, feel free to discuss it with us in the <a href="https://github.com/cityofsandiego/streetsSD/issues" rel="nofollow">issue queues</a>, fork the code, and make a Pull Request.</p>
]]></content:encoded>
      <guid>https://bymaksim.com/streetssd-overview</guid>
      <pubDate>Wed, 02 Nov 2016 01:32:05 +0000</pubDate>
    </item>
    <item>
      <title>What Is A Dataset? A Delicious Explanation</title>
      <link>https://bymaksim.com/what-is-a-dataset?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[If you work with data or interact with people who do, you have inevitably heard the word “dataset”. It’s one of those jargon-y words that gets thrown around, and people are somehow supposed to know what it means. It’s like when you go to the doctor, and he tells you that you will need “labrum repair”, and you feel stupid asking what a labrum is.&#xA;&#xA;!--more--&#xA;&#xA;Wikipedia’s Definition of a Dataset:&#xA;  Most commonly a data set corresponds to the contents of a single database table, or a single statistical data matrix, where every column of the table represents a particular variable, and each row corresponds to a given member of the data set in question.&#xA;&#xA;And data.gov’s definiton is:&#xA;  A dataset is an organized collection of data. The most basic representation of a dataset is data elements presented in tabular form. Each column represents a particular variable. Each row corresponds to a given value of that column’s variable. A dataset may also present information in a variety of non-tabular formats, such as an extended mark -up language (XML) file, a geospatial data file, or an image file.&#xA;&#xA;Nothing is incorrect about these definitions, but they tend to miscommunicate and under-explain what a dataset is. This causes a lot of confusion and frustration in the government open data community (and probably other communities as well). People end up thinking that there are datasets laying around on a shelf somewhere — that it’s a finite, well-defined thing that a city employee can just grab and publish. Sometimes, it’s true. However, more often than not, a dataset is more like this:&#xA;&#xA;There are two properties that every dataset has that are critical to how useful it is — format and structure.&#xA;&#xA;Format&#xA;&#xA;Format is fairly straightforward. In order for a dataset to be used by as many people as possible, it has to be open and machine-readable. This is just a primer — I’ll go into more detail in a different post.&#xA;&#xA;Open&#xA;&#xA;Open formats are those that can be read by a simple text editor. For example, CSV files — a common format for tabular data can be opened in Notepad. They will be hard to read by a human, but it’ll work. They can also be opened in Excel and plenty of languages have CSV parsers.&#xA;&#xA;However, the sister of the CSV file — the Excel file can only be opened in Excel. That means that if I want to look at your file, I have to give Microsoft some serious $$ for Microsoft Office.&#xA;&#xA;Machine Readable&#xA;&#xA;The other aspect of an open dataset is that it’s machine-readable. This ties in with “open” — since many more languages and libraries can read formats not copyrighted by a vendor. But this also means structured in a predictable manner. For example, a PDF file is not machine readable because it’s digital paper — and breaking down the structure of a PDF is really hard for a program that can only deal with structured data — the format is closed and copyrighted, and also extremely opaque.&#xA;&#xA;Structure&#xA;&#xA;Structure is a lot more difficult to get right than format. For the sake of this post (and to keep you awake), we’ll talk only about tabular data (data that lives in rows and columns). There are other types of data though, such as GIS (mapping) data, XML / JSON (nested data), and several others.&#xA;&#xA;Let’s step back and take a look at a database (or the dough blob). We’ll keep NoSQL databases out of this for simplicity and only focus on the old-school relational databases. What does a database look like?&#xA;&#xA;Contrary to what Hollywood would like us to think, this is not a database:&#xA;&#xA;A relational database isn’t really that exciting — it’s just a collection of tables (sorry, I tried, but relating this to cookies is hard):&#xA;&#xA;These tables contain a variety of columns. Invariably, one or more of these columns is going to have a set of values that relates to a set of values in another table. That’s what makes a database relational and also what gives the technology so much power and flexibility. It’s like dough — you can shape it, mold it, cut it and change its shape as much as you want to fit your cookie idea.&#xA;&#xA;Let’s take a look a more concrete example:&#xA;&#xA;Let’s say you have a company that sells widgets. You keep track of your customers, and you keep track of the widgets you sell. Let’s imagine a very simple database with just three tables:&#xA;&#xA;There is one table to keep information on your customers, another one to keep track of the widgets you have, and a table that keeps track of the orders — where the customers bought the widgets. For the sake of simplicity, there is only one customer per order and one widget per order.&#xA;&#xA;Let’s look at some example data in this type of structure and examine some points below:&#xA;&#xA;We can see that each table has an ID column. This is called a primary key and is used for the database to differentiate each record in that specific table. A lot of times it corresponds to a row number.&#xA;&#xA;In the orders table, we can see that each order has its own ID but also references the ID of the widget sold and the customer to whom it was sold.&#xA;&#xA;So with the following example, what are some datasets we can generate? There are many options:&#xA;&#xA;We can get 3 datasets already just by using the tables themselves. So, there will be a customers dataset, a widgets dataset, and an orders dataset. However, the orders dataset will be pretty useless, since it references IDs only relevant in the context of the other two tables.&#xA;&#xA;We can get a list of only those customers who have purchased a widget.&#xA;&#xA;We can get a list of customers who purchased a widget and live in Chicago.&#xA;&#xA;We can get a list of all the customers regardless of whether they have purchased a widget.&#xA;&#xA;We can get a list of customers, the widget they each purchased, the model of the widget, and how much the widget costs.&#xA;&#xA;We can get a list of widgets and which customers purchased them.&#xA;&#xA;Or a list of widgets and in which city they were purchased.&#xA;&#xA;We get these datasets by using SQL, a database query language, and what you see above is not an exhaustive list: there are many other alternatives. In addition, due to data schema changes, process changes and technology changes, data after a certain date may not be valid or may not mean the same thing. For example, up until our widget company changed their process in 2011, the “city” column in the customer table indicated the city in which the customer was born, not where he or she lived.&#xA;&#xA;This complexity is present in our simple fake database of just three tables, but modern relational databases have upwards of 100, often thousands of tables. There are also plenty of caveats in how they’re named, how relations are structured, and what type of queries make sense. Plus, writing the queries is no easy task.&#xA;&#xA;Some Clarification&#xA;&#xA;I want to take a moment to clarify here — I’m not saying that every dataset is hard to create. Sometimes it’s as easy as grabbing an Excel file from the desktop. But even then, it’s still necessary to make sure that there are no sensitive data in there, such as someone’s phone number.&#xA;&#xA;I also haven’t touched on things like aggregation and making sure that we generate data that fits the Tidy Data Spec so that it’s prime for analysis.&#xA;&#xA;How do you define a dataset?&#xA;&#xA;I hope I was able to show you that it’s not all cut and dry.&#xA;&#xA;The next time you go and ask someone for a dataset, bring them a box of cookies just in case. Fulfilling your request may involve a lot of hard work.&#xA;&#xA;Originally published at quandary.io on April 20, 2016.*&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>If you work with data or interact with people who do, you have inevitably heard the word “dataset”. It’s one of those jargon-y words that gets thrown around, and people are somehow supposed to know what it means. It’s like when you go to the doctor, and he tells you that you will need “labrum repair”, and you feel stupid asking what a labrum is.</p>



<p><a href="https://en.wikipedia.org/wiki/Data_set" rel="nofollow">Wikipedia’s Definition of a Dataset</a>:
&gt; <em>Most commonly a data set corresponds to the contents of a single database table, or a single statistical data matrix, where every column of the table represents a particular variable, and each row corresponds to a given member of the data set in question.</em></p>

<p>And <a href="http://www.data.gov/glossary" rel="nofollow">data.gov’s definiton</a> is:
&gt; <em>A dataset is an organized collection of data. The most basic representation of a dataset is data elements presented in tabular form. Each column represents a particular variable. Each row corresponds to a given value of that column’s variable. A dataset may also present information in a variety of non-tabular formats, such as an extended mark -up language (XML) file, a geospatial data file, or an image file.</em></p>

<p>Nothing is incorrect about these definitions, but they tend to miscommunicate and under-explain what a dataset is. This causes a lot of confusion and frustration in the government open data community (and probably other communities as well). People end up thinking that there are datasets laying around on a shelf somewhere — that it’s a finite, well-defined thing that a city employee can just grab and publish. Sometimes, it’s true. However, more often than not, a dataset is more like this:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*g3olp5pj7TgcO7qD." alt=""/></p>

<p>There are two properties that every dataset has that are critical to how useful it is — <strong>format</strong> and <strong>structure</strong>.</p>

<h2 id="format">Format</h2>

<p>Format is fairly straightforward. In order for a dataset to be used by as many people as possible, it has to be open and machine-readable. This is just a primer — I’ll go into more detail in a different post.</p>

<h3 id="open">Open</h3>

<p>Open formats are those that can be read by a simple text editor. For example, CSV files — a common format for tabular data can be opened in Notepad. They will be hard to read by a human, but it’ll work. They can also be opened in Excel and plenty of languages have CSV parsers.</p>

<p>However, the sister of the CSV file — the Excel file can <strong>only</strong> be opened in Excel. That means that if I want to look at your file, I have to give Microsoft some serious $$ for Microsoft Office.</p>

<h3 id="machine-readable">Machine Readable</h3>

<p>The other aspect of an open dataset is that it’s machine-readable. This ties in with “open” — since many more languages and libraries can read formats not copyrighted by a vendor. But this also means structured in a predictable manner. For example, a PDF file is not machine readable because it’s digital paper — and breaking down the structure of a PDF is really hard for a program that can only deal with structured data — the format is closed and copyrighted, and also extremely opaque.</p>

<h2 id="structure">Structure</h2>

<p>Structure is a lot more difficult to get right than format. For the sake of this post (and to keep you awake), we’ll talk only about tabular data (data that lives in rows and columns). There are other types of data though, such as GIS (mapping) data, XML / JSON (nested data), and several others.</p>

<p>Let’s step back and take a look at a database (or the dough blob). We’ll keep NoSQL databases out of this for simplicity and only focus on the old-school relational databases. What does a database look like?</p>

<p>Contrary to what Hollywood would like us to think, this is not a database:</p>

<p><img src="https://cdn-images-1.medium.com/max/2048/0*Osmg_NZvU2_tQyt0.jpg" alt=""/></p>

<p>A relational database isn’t really that exciting — it’s just a collection of tables (sorry, I tried, but relating this to cookies is hard):</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*Q2pgTYDnFO6Iy8kr." alt=""/></p>

<p>These tables contain a variety of columns. Invariably, one or more of these columns is going to have a set of values that relates to a set of values in another table. That’s what makes a database relational and also what gives the technology so much power and flexibility. It’s like dough — you can shape it, mold it, cut it and change its shape as much as you want to fit your cookie idea.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*5Bo1yYHHKdLGmgq8." alt=""/></p>

<p>Let’s take a look a more concrete example:</p>

<p>Let’s say you have a company that sells widgets. You keep track of your customers, and you keep track of the widgets you sell. Let’s imagine a very simple database with just three tables:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*qhH_XKZ22NRVt6nk." alt=""/></p>

<p>There is one table to keep information on your customers, another one to keep track of the widgets you have, and a table that keeps track of the orders — where the customers bought the widgets. For the sake of simplicity, there is only one customer per order and one widget per order.</p>

<p>Let’s look at some example data in this type of structure and examine some points below:</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*4N3w4asMfK5HHGGR." alt=""/></p>
<ul><li><p>We can see that each table has an ID column. This is called a <strong>primary key</strong> and is used for the database to differentiate each record in that specific table. A lot of times it corresponds to a row number.</p></li>

<li><p>In the orders table, we can see that each order has its own ID but also references the ID of the widget sold and the customer to whom it was sold.</p></li></ul>

<p>So with the following example, what are some datasets we can generate? There are many options:</p>
<ul><li><p>We can get 3 datasets already just by using the tables themselves. So, there will be a customers dataset, a widgets dataset, and an orders dataset. However, the orders dataset will be pretty useless, since it references IDs only relevant in the context of the other two tables.</p></li>

<li><p>We can get a list of only those customers who have purchased a widget.</p></li>

<li><p>We can get a list of customers who purchased a widget and live in Chicago.</p></li>

<li><p>We can get a list of all the customers regardless of whether they have purchased a widget.</p></li>

<li><p>We can get a list of customers, the widget they each purchased, the model of the widget, and how much the widget costs.</p></li>

<li><p>We can get a list of widgets and which customers purchased them.</p></li>

<li><p>Or a list of widgets and in which city they were purchased.</p></li></ul>

<p>We get these datasets by using SQL, a database query language, and what you see above is not an exhaustive list: there are many other alternatives. In addition, due to data schema changes, process changes and technology changes, data after a certain date may not be valid or may not mean the same thing. For example, up until our widget company changed their process in 2011, the “city” column in the customer table indicated the city in which the customer was born, not where he or she lived.</p>

<p>This complexity is present in our simple fake database of just three tables, but modern relational databases have upwards of 100, often thousands of tables. There are also plenty of caveats in how they’re named, how relations are structured, and what type of queries make sense. Plus, writing the queries is no easy task.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*x7k_vrissUquID8Q." alt=""/></p>

<h2 id="some-clarification">Some Clarification</h2>

<p>I want to take a moment to clarify here — I’m not saying that every dataset is hard to create. Sometimes it’s as easy as grabbing an Excel file from the desktop. But even then, it’s still necessary to make sure that there are no sensitive data in there, such as someone’s phone number.</p>

<p>I also haven’t touched on things like aggregation and making sure that we generate data that fits the <a href="http://vita.had.co.nz/papers/tidy-data.pdf" rel="nofollow">Tidy Data Spec</a> so that it’s prime for analysis.</p>

<p>How do <strong><em>you</em></strong> define a dataset?</p>

<p>I hope I was able to show you that it’s not all cut and dry.</p>

<p>The next time you go and ask someone for a dataset, bring them a box of cookies just in case. Fulfilling your request may involve a lot of hard work.</p>

<p><em>Originally published at <a href="http://quandary.io/what-is-a-dataset/" rel="nofollow">quandary.io</a> on April 20, 2016.</em></p>
]]></content:encoded>
      <guid>https://bymaksim.com/what-is-a-dataset</guid>
      <pubDate>Tue, 19 Apr 2016 21:19:47 +0000</pubDate>
    </item>
    <item>
      <title>What I Learned In One Year as CDO of San Diego</title>
      <link>https://bymaksim.com/what-i-learned-in-one-year-as-cdo-of-san-diego?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[&#xA;&#xA;Fairly recently, I celebrated my 1 year anniversary of being with the city of San Diego as the Chief Data Officer. Naturally, I’ve been reflecting on my first year in government, and one of the things that strikes me the most is the amount of things that I learned.&#xA;&#xA;!--more--&#xA;&#xA;Some things I learned on my own — in many cases I invested in supplementary learning in order to succeed. Others I can directly attribute to conversations with my boss, Almis, my teammates in Performance and Analytics, the City’s IT staff, and so many others that if I list them all, this would be a spreadsheet, not a blog post.&#xA;&#xA;To be completely honest — I had my hesitations about taking a local government job. I was worried I wouldn’t advance my technological skill, and potentially lose many of the marketable skills I possess. However, looking back on this year, taking a government job has proven to be one of the smartest and challenging things I’ve ever done. There is no other place I can think of that would push me so far out of my comfort zone, in so many different directions, while providing me the help and support to succeed.&#xA;&#xA;Technical Skills&#xA;&#xA;Tools for the Job. I’m proud to say that I have a very wide expertise in a diverse range of technologies, languages and stacks. I walked into this job with a preconception that my large toolbox can solve a variety of problems quickly. Instead, I quickly realized that in order to use my toolbox effectively, I needed to understand all the intricate non-technological complexities such as strategy, culture, processes, and even words used to describe things. What’s more, I had to expand my toolbox with more open source and not-so-open-source tools. The City is a large enterprise, and not everything can be solved by cloning a repo or a $30 a month subscription. This intangible skill of understanding City operations and figuring out the details of how to apply the correct technology has been one of my biggest learnings this year.&#xA;&#xA;R Language. This past year, I learned the R language — the first data-focused language I’ve used. In addition to learning the language itself, I learned how to communicate research and data analysis through detailed and reproducible reports.&#xA;&#xA;Google Apps Script. I wouldn’t necessarily say that I “learned” it, but I definitely became much more familiar with the capacities of Apps Script for handling large amounts of documents programmatically. This code ran a lot of our initial data inventory.&#xA;&#xA;GIS. Through our wonderful GIS manager, ArcGIS and CartoDB, I became much more familiar with GIS operations and concepts. However, I would place myself very far from expert level.&#xA;&#xA;Cybersecurity. Something we often don’t think about when it comes to cities is the massive need for cybersecurity. Cities are constantly getting hit and need robust and well-managed cybersecurity programs in place. I have become good friends with the CISO and the Cyber team in San Diego and learned quite a bit about their strategies, technologies and methods.&#xA;&#xA;Enterprise IT. We often like to make fun of “antiquated” City technology like SAP and Oracle. However, these are still massive companies running multi-million dollar projects and are not going away soon. I’ve worked quite a bit with IT to understand the ecosystems and infrastructure that manage these large datastores within the City.&#xA;&#xA;So Many Datastores. This brings me to the diversity of datastores within the City. From a technical level, I’ve been exposed to a variety of SAP, Oracle, Microsoft and ESRI products. This year, I’ve had lots of technology land on my plate. Even though this isn’t very sexy, going down the journey of learning these technologies has been a blast.&#xA;&#xA;Networks. Working with IT to understand the City’s networks, I’ve learned a ton about the City’s DNS architecture, networks, configuration management, desktop management, Active Directory and so much more. The City also does a lot of cool things, such as producing its own TV station. We even have a truck that can pull up to a location and provide wi-fi and cellular communications in emergencies.&#xA;&#xA;Strategic Skills&#xA;&#xA;Technological Interventions. Operational improvement opportunities that would otherwise remain hidden often surface when a public-facing data visualization or analysis is on the horizon. It’s a fortunate thing I work in a department that cares to fix those as well. I have also begun to explore how to use what I call “technological interventions” strategically to discover problems in processes of collecting, generating and sharing data across systems.&#xA;&#xA;Communicating with Executives. In the course of my work, I had to learn how to work with elected officials and high level executives. I learned how to properly scale the level of detail while explaining complex technical concepts and challenges. Scaling detail is extremely important because it’s often necessary for me to walk a fine line between confusing the person I’m talking to and providing them with the amount of information necessary for them to make a decision.&#xA;&#xA;Managing up became very important as well. My boss is extremely flexible and understanding, but oftentimes he doesn’t need to (or really want to) be deeply involved in the details of my decision making and thought processes: he empowers me, which is both scary and great. It makes sense too: he really doesn’t need to know all the variables I’m weighing when deciding on a technology or working on a project. By managing up, I’m able to keep his expectations at the proper levels while giving myself a good amount of flexibility to change my mind.&#xA;&#xA;Managing Meetings. Government is notorious for having a meeting-loving culture. I learned how to manage and limit the amount of meetings that I end up involved with, while making sure that I maintain relationships and remain apprised of what’s going on.&#xA;&#xA;Evaluating technology in the context of the City, or rather in the context of enterprise IT requires weighing many factors. I learned a lot about how to evaluate technological solutions in this type of context. How does a piece of software fit into the larger IT roadmap? What do the service level agreements and maintenance options look like? Is it better to get Open Source or proprietary solutions? How will this impact the City’s cybersecurity risk exposure? What is the price point and what are the related procurement rules dictating the process we’ll go through? All these factors weigh into evaluating a piece of software in a large IT environment.&#xA;&#xA;Hacking People, no USB cable required. No, really, I learned how to get 65 non-technical people who have other jobs to answer my questions about what data they have, make it extremely easy for them to do so, and not have them hate you afterwards. Oh, and only have one large meeting through that whole process.&#xA;&#xA;Shipping. I learned quite a bit about how to balance the inherent complexity of being a CDO of a massive organization with designing the correct architecture, building a sustainable infrastructure, and well, shipping.&#xA;&#xA;More Shipping. Elaborating on the point above, I tend to lean more on the perfectionist side of being a programmer. I’ll think deeply about choosing the right framework, picking the right stack and overall making sure I’m doing things “right”. In the past, this caused unnecessary slowdowns in some of my projects. However, in this past year, when I had a much smaller part of my time to be a developer, I noticed myself being able to adjust my “perfectionist” strategies towards a strategy where I just write some code, and refactor if I’m going to keep it. It seems like a minor shift, but being able to accept imperfection in my own code has allowed me to prototype a lot more, a lot faster.&#xA;&#xA;Technology Perception Gaps. I’ve learned to see and more clearly understand the gap between the public’s perception of Open Data and the technology and effort needed to deliver data that isn’t just open, but is accurate and complete. While this gap exists in the Open Data space, I would argue that this gap exists in many perceptions that the public has about City IT.&#xA;&#xA;Learning from Private Sector. I often seek out sources of inspiration for my work by attending private sector events, training sessions, and simply communicating with technology communities that I’m used to working with. I also love to interact with other governments to see how they’re doing things. However, not everything that works in the private sector works in the government space. I learned quite a bit about how to learn (and not learn) from the private sector and relate it to our work at the City.&#xA;&#xA;Open Data Activation Points. I’ve begun to see “Activation Points” for open data — where opening data ties in smooth as butter with a current City initiative. (Sustainability, Performance Management, Public Records Act Requests (PRA)).&#xA;&#xA;Building and running a data team is a leadership, strategic and tactical challenge. I have learned, and am still learning a lot about how to build one, ramp it up, and align the correct pieces for delivery.&#xA;&#xA;Departments / City Operations&#xA;&#xA;Lifeguards. I learned a ton about what lifeguards do, how they operate, and some of the technological challenges of the job.&#xA;&#xA;Police. I explored how the City’s Police Department works with bar owners in the Gaslamp to create a safer area, handle barbreak (when all the bars let out) and prevent crime.&#xA;&#xA;Street Work. I know much more about how the City’s street work gets done, and the different types of repair jobs. I was also surprised to find out that the City does a proactive job of making sure that money isn’t wasted by having the same street repaved multiple times in a short period. Because of various responsibilities different departments have when it comes to digging up a street, they do a lot of work to coordinate the timing of the construction.&#xA;&#xA;PRAs. Many people outside of government or journalism don’t even know about PRAs. Working with the City’s PRA coordinator, I’ve learned a ton about the legalities of the process and how they work within the City.&#xA;&#xA;Emergency Response. Do you know what to do if an earthquake hits? Your City officials do — they often conduct training and simulations for different disaster scenarios. I got to explore this as well. They have a “situation room” in one of the buildings, stocked with food and computers. Learning about the City’s ways of managing disaster has been nothing short of amazing.&#xA;&#xA;Crap. I learned all about what happens to the crap we flush. It’s actually way more interesting than it sounds. This is a story that encompasses poop, energy generation, earplugs, and seals. Where did I learn about this? The Point Loma Wastewater Treatment Plant — an automated plant that serves 2.2 million residents run by only 50 people.&#xA;&#xA;Communications and Community&#xA;&#xA;Talking Louder. I have been a quiet and shy speaker for most of my life. However, when I started this job, I quickly learned (with Almis’ advice) that I needed to overcome that. I took several months of improv classes (which were a blast) and while I’m still working up to having a booming voice, I can definitely say that I haven’t been asked to speak louder in a long time.&#xA;&#xA;Press. I learned a lot about how to talk and work with the press. It was really fascinating to learn the role of press in government and how the City handles communication with journalists.&#xA;&#xA;Public Speaking. I still need to tally up how much public speaking I’ve done this year, but it was definitely more than I’ve ever done in my life. There’s plenty of things I learned in this arena this year, such as structuring effective presentations, storytelling, and other skills.&#xA;&#xA;Ignite Talk. I gave an Ignite Talk about the data flowing through a City block. Obviously, I learned how to give an Ignite talk, which in itself is a very difficult, but ultimately rewarding experience. But through that process I discovered how much I knew about the data flowing through cities — and how much I have yet to learn.&#xA;&#xA;Reporting to Council. I also got to prepare a report to the City council and present it — a learning experience in and of itself. As a side note, as part of getting the Open Data Implementation Update ready for council, I ended up making a pull request to a major open-source software project.&#xA;&#xA;Community Building. I always felt that in order for Open Data to be successful in San Diego, we needed to have a strong Civic Tech community. Before, this was far from the case. Now, we’re much closer — I’ve learned a ton about how to put the right people in the right place to get the community moving. I’m very excited about the work Open San Diego has accomplished in the past year.&#xA;&#xA;Connecting Geeks to Government. In addition to building the community, I’ve always strongly felt there was a missing connection between CfA Brigade members, residents, and City employees. Bringing various people in the City-people I greatly respect-to talk about what they do with the brigade has been an invaluable learning experience both for me, the brigade, and the employees. I think we often forget how freaking cool City government can be.&#xA;&#xA;Surprises&#xA;&#xA;Passion. I was surprised about how I feel after a conversation with a City employee (not everyone of course, but so many of them). For example, I talked to someone about parking meters and was inspired and excited after that conversation. It’s amazing to see people that have just as much passion for their craft as many programmers do for theirs.&#xA;&#xA;Data geeks. I’m not talking “analysts”. I’m talking about the people that geek out over Excel macros and SQL queries. Or GIS. There aren’t many of them in the City, but they’re there. It takes a while to find them too. But once you do — wow — they can teach you so much and are such a pleasure to work with.&#xA;&#xA;Data Literacy. How people think about what data is, how they define data, how they understand KPIs and metrics has been a challenge that I have enjoyed learning about and solving.&#xA;&#xA;Open Doors. Many people already recognize the potential for decreased workload and better communications that opening data carries. I was extremely surprised how many people approached me, excited for their data to be opened up.&#xA;&#xA;State of The City. Did you know that just like the President does an annual State of the Union, many mayors do a State of the City Address? I didn’t. It was amazing to see it get put together, how many people work on it, and some of the things I learned by attending. Also it was pretty sweet to get a shout from Mayor Faulconer just a month into the job.&#xA;&#xA;Dependencies. I noticed a distinct difference between how I’m used to working to how I have to work in the City. In my previous jobs as a software developer and architect, I could generally forge my own path and minimize a lot of critical reliance on other people. In the City I feel highly leveraged — oftentimes the support of others is absolutely critical. I haven’t thought too far down this path yet, but will be exploring it.&#xA;&#xA;Internationalizing. I’ve had the opportunity to present to Macedonian and Philippine delegations. In the Mayor’s conference room. About Open Data. Nuff said.&#xA;&#xA;CNAMEs. When we were launching OpenGov — our budget tool, I thought it would be great for us to map it to http://budget.sandiego.gov. The saga that happened while I was trying to do something fairly simple — just map a CNAME to a domain — was a journey I will never forget.&#xA;&#xA;Random&#xA;&#xA;Vendors. Working with vendors has been good, bad and ugly. There is so much to write about here: from evaluating proposals and negotiating; to the “classifications” of vendors I’ve begun to unofficially define based on what they sell and how they sell it; to how they can have better strategy and products. We have developed numerous ways to streamline the process of vendor pitches by having a standard method and questions that they have to use to get in touch with us.&#xA;&#xA;City Awards and Competitions, run by various think tanks and vendors are also something I’ve been reflecting on lately. While having some value, a lot of times they appear to be run more for the benefit of a vendor looking to score more government contracting opportunities.&#xA;&#xA;Time Management. I’ve had to implement rigid time management methodologies because of the amount of noise often coming on my plate. While this may seem like an underwhelming learned lesson, I would say this is what has allowed me to become the most effective.&#xA;&#xA;Project Management. When you have an organization of 11,000 people, IT projects quickly start to look a little different. Sometimes there’s a legitimate need for a strong project management methodology. Sometimes there isn’t. I’ve learned quite a bit about how the City runs large IT projects–and small ones.&#xA;&#xA;Hiring. As part of hiring my first employee, I got to be on the hiring side, all the way through the recruitment process. I ended up taking notes for myself, for, you know, the next time I have to look for a job. It was super valuable.&#xA;&#xA;Supervising. Because I was getting a new employee, I had to go through the City’s “Supervisor’s Academy”. This was a seven day course, outlining everything from City support for employees experiencing various issues, to ethics, to leadership. I found it extremely useful and learned a ton about, well, being a good supervisor.&#xA;&#xA;Ethics. I learned a bit of how ethics and campaign finance work in the City context, and had some fairly interesting conversations with the Ethics commission, including one about the limits on what expenses can and cannot be covered by a conference organizer who invites me to present at a conference.&#xA;&#xA;In summarizing my experience at the City, the main thing that surprised me this year was how much I was pushed out of my boundaries, and how many things I’ve done that I never thought I could or would do. Everything from learning a new language, doing an Ignite talk, to learning how to be a better communicator and leader, I can attribute to being a City employee in one way or another. Now that year 1 is behind me, I’m stoked for year 2.&#xA;&#xA;Originally published at mrmaksimize.com.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p><img src="https://cdn-images-1.medium.com/max/2800/0*pZyXB3L3u4VDVYm0.jpeg" alt=""/></p>

<p>Fairly recently, I celebrated my 1 year anniversary of being with the city of San Diego as the Chief Data Officer. Naturally, I’ve been reflecting on my first year in government, and one of the things that strikes me the most is the amount of things that I learned.</p>



<p>Some things I learned on my own — in many cases I invested in supplementary learning in order to succeed. Others I can directly attribute to conversations with my boss, Almis, my teammates in <a href="http://sandiego.gov/pad" rel="nofollow">Performance and Analytics</a>, the City’s IT staff, and so many others that if I list them all, this would be a spreadsheet, not a blog post.</p>

<p>To be completely honest — I had my hesitations about taking a local government job. I was worried I wouldn’t advance my technological skill, and potentially lose many of the marketable skills I possess. However, looking back on this year, taking a government job has proven to be one of the smartest and challenging things I’ve ever done. There is no other place I can think of that would push me so far out of my comfort zone, in so many different directions, while providing me the help and support to succeed.</p>

<h2 id="technical-skills">Technical Skills</h2>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*CnkjBdu74QGfgqbx.jpg" alt=""/></p>

<p><strong>Tools for the Job.</strong> I’m proud to say that I have a very wide expertise in a diverse range of technologies, languages and stacks. I walked into this job with a preconception that my large toolbox can solve a variety of problems quickly. Instead, I quickly realized that in order to use my toolbox effectively, I needed to understand all the intricate non-technological complexities such as strategy, culture, processes, and even words used to describe things. What’s more, I had to expand my toolbox with more open source and not-so-open-source tools. The City is a large enterprise, and not everything can be solved by cloning a repo or a $30 a month subscription. This intangible skill of understanding City operations and figuring out the details of how to apply the correct technology has been one of my biggest learnings this year.</p>

<p><strong>R Language.</strong> This past year, I learned the <a href="https://www.r-project.org/" rel="nofollow">R language</a> — the first data-focused language I’ve used. In addition to learning the language itself, I learned how to communicate research and data analysis through detailed and reproducible reports.</p>

<p><strong>Google Apps Script.</strong> I wouldn’t necessarily say that I “learned” it, but I definitely became much more familiar with the capacities of Apps Script for handling large amounts of documents programmatically. This code ran a lot of our initial data inventory.</p>

<p><strong>GIS.</strong> Through our wonderful GIS manager, <a href="https://www.arcgis.com/features/" rel="nofollow">ArcGIS</a> and <a href="https://cartodb.com/" rel="nofollow">CartoDB</a>, I became much more familiar with GIS operations and concepts. However, I would place myself very far from expert level.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*3Pl3RGWbGC_z2cCJ.png" alt=""/></p>

<p><strong>Cybersecurity.</strong> Something we often don’t think about when it comes to cities is the massive need for cybersecurity. Cities are constantly getting hit and need robust and well-managed cybersecurity programs in place. I have become good friends with the CISO and the Cyber team in San Diego and learned quite a bit about their strategies, technologies and methods.</p>

<p><strong>Enterprise IT.</strong> We often like to make fun of “antiquated” City technology like SAP and Oracle. However, these are still massive companies running multi-million dollar projects and are not going away soon. I’ve worked quite a bit with IT to understand the ecosystems and infrastructure that manage these large datastores within the City.</p>

<p><strong>So Many Datastores.</strong> This brings me to the diversity of datastores within the City. From a technical level, I’ve been exposed to a variety of SAP, Oracle, Microsoft and ESRI products. This year, I’ve had lots of technology land on my plate. Even though this isn’t very sexy, going down the journey of learning these technologies has been a blast.</p>

<p><strong>Networks.</strong> Working with IT to understand the City’s networks, I’ve learned a ton about the City’s DNS architecture, networks, configuration management, desktop management, Active Directory and so much more. The City also does a lot of cool things, such as producing its own TV station. We even have a truck that can pull up to a location and provide wi-fi and cellular communications in emergencies.</p>

<h2 id="strategic-skills">Strategic Skills</h2>

<p><strong>Technological Interventions.</strong> Operational improvement opportunities that would otherwise remain hidden often surface when a public-facing data visualization or analysis is on the horizon. It’s a fortunate thing I work in a department that cares to fix those as well. I have also begun to explore how to use what I call “technological interventions” strategically to discover problems in processes of collecting, generating and sharing data across systems.</p>

<p><strong>Communicating with Executives.</strong> In the course of my work, I had to learn how to work with elected officials and high level executives. I learned how to properly scale the level of detail while explaining complex technical concepts and challenges. Scaling detail is extremely important because it’s often necessary for me to walk a fine line between confusing the person I’m talking to and providing them with the amount of information necessary for them to make a decision.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*RLKK5rMUBIQuoilJ.png" alt=""/></p>

<p><strong>Managing up</strong> became very important as well. My boss is extremely flexible and understanding, but oftentimes he doesn’t need to (or really want to) be deeply involved in the details of my decision making and thought processes: he empowers me, which is both scary and great. It makes sense too: he really doesn’t need to know all the variables I’m weighing when deciding on a technology or working on a project. By managing up, I’m able to keep his expectations at the proper levels while giving myself a good amount of flexibility to change my mind.</p>

<p><strong>Managing Meetings.</strong> Government is notorious for having a meeting-loving culture. I learned how to manage and limit the amount of meetings that I end up involved with, while making sure that I maintain relationships and remain apprised of what’s going on.</p>

<p><strong>Evaluating technology</strong> in the context of the City, or rather in the context of enterprise IT requires weighing many factors. I learned a lot about how to evaluate technological solutions in this type of context. How does a piece of software fit into the larger IT roadmap? What do the service level agreements and maintenance options look like? Is it better to get Open Source or proprietary solutions? How will this impact the City’s cybersecurity risk exposure? What is the price point and what are the related procurement rules dictating the process we’ll go through? All these factors weigh into evaluating a piece of software in a large IT environment.</p>

<p><strong>Hacking People, no USB cable required</strong>. No, really, I learned how to get 65 non-technical people who have other jobs to answer my questions about what data they have, make it extremely easy for them to do so, and not have them hate you afterwards. Oh, and only have one large meeting through that whole process.</p>

<p><strong>Shipping.</strong> I learned quite a bit about how to balance the inherent complexity of being a CDO of a massive organization with designing the correct architecture, building a sustainable infrastructure, and well, shipping.</p>

<p><strong>More Shipping.</strong> Elaborating on the point above, I tend to lean more on the perfectionist side of being a programmer. I’ll think deeply about choosing the right framework, picking the right stack and overall making sure I’m doing things “right”. In the past, this caused unnecessary slowdowns in some of my projects. However, in this past year, when I had a much smaller part of my time to be a developer, I noticed myself being able to adjust my “perfectionist” strategies towards a strategy where I just write some code, and refactor if I’m going to keep it. It seems like a minor shift, but being able to accept imperfection in my own code has allowed me to prototype a lot more, a lot faster.</p>

<p><strong>Technology Perception Gaps.</strong> I’ve learned to see and more clearly understand the gap between the public’s perception of Open Data and the technology and effort needed to deliver data that isn’t just open, but is accurate and complete. While this gap exists in the Open Data space, I would argue that this gap exists in many perceptions that the public has about City IT.</p>

<p><strong>Learning from Private Sector.</strong> I often seek out sources of inspiration for my work by attending private sector events, training sessions, and simply communicating with technology communities that I’m used to working with. I also love to interact with other governments to see how they’re doing things. However, not everything that works in the private sector works in the government space. I learned quite a bit about how to learn (and not learn) from the private sector and relate it to our work at the City.</p>

<p><strong>Open Data Activation Points.</strong> I’ve begun to see “Activation Points” for open data — where opening data ties in smooth as butter with a current City initiative. (Sustainability, Performance Management, Public Records Act Requests (PRA)).</p>

<p><strong>Building and running a data team</strong> is a leadership, strategic and tactical challenge. I have learned, and am still learning a lot about how to build one, ramp it up, and align the correct pieces for delivery.</p>

<h2 id="departments-city-operations">Departments / City Operations</h2>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*yE7Lei3F9ihCjezk.jpg" alt=""/></p>

<p><strong>Lifeguards.</strong> I learned a ton about what lifeguards do, how they operate, and some of the technological challenges of the job.</p>

<p><strong>Police.</strong> I explored how the City’s Police Department works with bar owners in the Gaslamp to create a safer area, handle barbreak (when all the bars let out) and prevent crime.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*SWOpjhaNOqtGgBVv.jpg" alt=""/></p>

<p><strong>Street Work.</strong> I know much more about how the City’s street work gets done, and the different types of repair jobs. I was also surprised to find out that the City does a proactive job of making sure that money isn’t wasted by having the same street repaved multiple times in a short period. Because of various responsibilities different departments have when it comes to digging up a street, they do a lot of work to coordinate the timing of the construction.</p>

<p><strong>PRAs.</strong> Many people outside of government or journalism don’t even know about PRAs. Working with the City’s PRA coordinator, I’ve learned a ton about the legalities of the process and how they work within the City.</p>

<p><strong>Emergency Response.</strong> Do you know what to do if an earthquake hits? Your City officials do — they often conduct training and simulations for different disaster scenarios. I got to explore this as well. They have a “situation room” in one of the buildings, stocked with food and computers. Learning about the City’s ways of managing disaster has been nothing short of amazing.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*m06j1-P4L__-G_ti.jpg" alt=""/></p>

<p><strong>Crap.</strong> I learned all about what happens to the crap we flush. It’s actually way more interesting than it sounds. This is a story that encompasses poop, energy generation, earplugs, and seals. Where did I learn about this? The <a href="http://www.sandiego.gov/mwwd/facilities/ptloma/" rel="nofollow">Point Loma Wastewater Treatment Plant</a> — an automated plant that serves 2.2 million residents run by only 50 people.</p>

<h2 id="communications-and-community">Communications and Community</h2>

<p><strong>Talking Louder.</strong> I have been a quiet and shy speaker for most of my life. However, when I started this job, I quickly learned (with Almis’ advice) that I needed to overcome that. I took several months of improv classes (which were a blast) and while I’m still working up to having a booming voice, I can definitely say that I haven’t been asked to speak louder in a long time.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*qtcCyjsLsYyLqINU.jpg" alt=""/></p>

<p><strong>Press.</strong> I learned a lot about how to talk and work with the press. It was really fascinating to learn the role of press in government and how the City handles communication with journalists.</p>

<p><strong>Public Speaking.</strong> I still need to tally up how much public speaking I’ve done this year, but it was definitely more than I’ve ever done in my life. There’s plenty of things I learned in this arena this year, such as structuring effective presentations, storytelling, and other skills.</p>

<p><img src="https://cdn-images-1.medium.com/max/2000/0*WeqpBIv53XJwIJts.png" alt=""/></p>

<p><strong>Ignite Talk.</strong> I gave an <a href="https://www.youtube.com/watch?v=IvABvAM11XM" rel="nofollow">Ignite Talk</a> about the data flowing through a City block. Obviously, I learned how to give an Ignite talk, which in itself is a very difficult, but ultimately rewarding experience. But through that process I discovered how much I knew about the data flowing through cities — and how much I have yet to learn.</p>

<p><strong>Reporting to Council.</strong> I also got to prepare a report to the City council and present it — a learning experience in and of itself. As a side note, as part of getting the <a href="https://datasd.gitbooks.io/council_report/" rel="nofollow">Open Data Implementation Update</a> ready for council, I ended up making a <a href="https://github.com/GitbookIO/gitbook/pull/813" rel="nofollow">pull request</a> to a major open-source software project.</p>

<p><strong>Community Building.</strong> I always felt that in order for Open Data to be successful in San Diego, we needed to have a strong Civic Tech community. Before, this was far from the case. Now, we’re much closer — I’ve learned a ton about how to put the right people in the right place to get the community moving. I’m very excited about the <a href="https://github.com/opensandiego" rel="nofollow">work</a> <a href="http://www.meetup.com/Open-San-Diego/" rel="nofollow">Open San Diego</a> has accomplished in the past year.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*KexIECs1MqkliwrW.jpg" alt=""/></p>

<p><strong>Connecting Geeks to Government.</strong> In addition to building the community, I’ve always strongly felt there was a missing connection between CfA Brigade members, residents, and City employees. Bringing various people in the City-people I greatly respect-to talk about what they do with the brigade has been an invaluable learning experience both for me, the brigade, and the employees. I think we often forget how freaking cool City government can be.</p>

<h2 id="surprises">Surprises</h2>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*h7dDif6_PD_nsJoU.jpg" alt=""/></p>

<p><strong>Passion.</strong> I was surprised about how I feel after a conversation with a City employee (not everyone of course, but so many of them). For example, I talked to someone about parking meters and was inspired and excited after that conversation. It’s amazing to see people that have just as much passion for their craft as many programmers do for theirs.</p>

<p><strong>Data geeks.</strong> I’m not talking “analysts”. I’m talking about the people that geek out over Excel macros and SQL queries. Or GIS. There aren’t many of them in the City, but they’re there. It takes a while to find them too. But once you do — wow — they can teach you so much and are such a pleasure to work with.</p>

<p><strong>Data Literacy.</strong> How people think about what data is, how they define data, how they understand KPIs and metrics has been a challenge that I have enjoyed learning about and solving.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*7U6pnsmE7CtxH6RT.jpg" alt=""/></p>

<p><strong>Open Doors.</strong> Many people already recognize the potential for decreased workload and better communications that opening data carries. I was extremely surprised how many people approached me, excited for their data to be opened up.</p>

<p><strong>State of The City.</strong> Did you know that just like the President does an annual State of the Union, many mayors do a State of the City Address? I didn’t. It was amazing to see it get put together, how many people work on it, and some of the things I learned by attending. Also it was pretty sweet to get a shout from Mayor Faulconer just a month into the job.</p>

<p><strong>Dependencies.</strong> I noticed a distinct difference between how I’m used to working to how I have to work in the City. In my previous jobs as a software developer and architect, I could generally forge my own path and minimize a lot of critical reliance on other people. In the City I feel highly leveraged — oftentimes the support of others is absolutely critical. I haven’t thought too far down this path yet, but will be exploring it.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*W4CpVZn1zHjoq5rz.jpg" alt=""/></p>

<p><strong>Internationalizing.</strong> I’ve had the opportunity to present to Macedonian and Philippine delegations. In the Mayor’s conference room. About Open Data. Nuff said.</p>

<p><strong>CNAMEs.</strong> When we were launching OpenGov — our budget tool, I thought it would be great for us to map it to <a href="http://budget.sandiego.gov" rel="nofollow">http://budget.sandiego.gov</a>. The saga that happened while I was trying to do something fairly simple — just map a CNAME to a domain — was a journey I will never forget.</p>

<h2 id="random">Random</h2>

<p><strong>Vendors.</strong> Working with vendors has been good, bad and ugly. There is so much to write about here: from evaluating proposals and negotiating; to the “classifications” of vendors I’ve begun to unofficially define based on what they sell and how they sell it; to how they can have better strategy and products. We have developed numerous ways to streamline the process of vendor pitches by having a standard method and questions that they have to use to get in touch with us.</p>

<p><strong>City Awards and Competitions</strong>, run by various think tanks and vendors are also something I’ve been reflecting on lately. While having some value, a lot of times they appear to be run more for the benefit of a vendor looking to score more government contracting opportunities.</p>

<p><strong>Time Management.</strong> I’ve had to implement rigid time management methodologies because of the amount of noise often coming on my plate. While this may seem like an underwhelming learned lesson, I would say this is what has allowed me to become the most effective.</p>

<p><strong>Project Management.</strong> When you have an organization of 11,000 people, IT projects quickly start to look a little different. Sometimes there’s a legitimate need for a strong project management methodology. Sometimes there isn’t. I’ve learned quite a bit about how the City runs large IT projects–and small ones.</p>

<p><strong>Hiring.</strong> As part of hiring my first employee, I got to be on the hiring side, all the way through the recruitment process. I ended up taking notes for myself, for, you know, the next time I have to look for a job. It was super valuable.</p>

<p><img src="https://cdn-images-1.medium.com/max/2800/0*vJbtHHcLbtAsm8r-.jpg" alt=""/></p>

<p><strong>Supervising.</strong> Because I was getting a new employee, I had to go through the City’s “Supervisor’s Academy”. This was a seven day course, outlining everything from City support for employees experiencing various issues, to ethics, to leadership. I found it extremely useful and learned a ton about, well, being a good supervisor.</p>

<p><strong>Ethics.</strong> I learned a bit of how ethics and campaign finance work in the City context, and had some fairly interesting conversations with the Ethics commission, including one about the limits on what expenses can and cannot be covered by a conference organizer who invites me to present at a conference.</p>

<p>In summarizing my experience at the City, the main thing that surprised me this year was how much I was pushed out of my boundaries, and how many things I’ve done that I never thought I could or would do. Everything from learning a new language, doing an Ignite talk, to learning how to be a better communicator and leader, I can attribute to being a City employee in one way or another. Now that year 1 is behind me, I’m stoked for year 2.</p>

<p><em>Originally published at <a href="http://mrmaksimize.com/what-i-learned-one-year-cdo/" rel="nofollow">mrmaksimize.com</a>.</em></p>
]]></content:encoded>
      <guid>https://bymaksim.com/what-i-learned-in-one-year-as-cdo-of-san-diego</guid>
      <pubDate>Mon, 11 Jan 2016 13:07:37 +0000</pubDate>
    </item>
  </channel>
</rss>