top of page

Proptech Spotlight – Transforming Fragmented Data Into AI-ready Assets

  • Jul 29
  • 37 min read

In the race to adopt AI, many proptech organisations are discovering a hard truth – their data is not ready. Fragmented, inconsistent and incomplete datasets are holding back the very innovations that promise to transform the property industry.


That was the focus of a recent Proptech Spotlight webinar, hosted by Proptech Australia President Kylie Davis last Tuesday, 21 July. The free online session brought together two industry experts to tackle one of the sector's most pressing challenges: how to turn messy, siloed data into a trusted foundation for AI.


Speakers:

  • Gerry Stanley (Precisely)

  • Michael Ogilvie (Data Army)

    Hosted By: Kylie Davis, President of Proptech Australia


Together, the speakers walked through practical examples and real-world use cases demonstrating how organisations have improved data quality, accelerated AI adoption and achieved measurable business outcomes.


Transcript:


Kylie: Hello, everyone. If you're joining us, it's Kylie Davis from Proptech Australia here. It's great to see so many people jumping on the call. We're here to talk about transforming fragmented data into AI-ready assets, and we're just gonna give everyone a few minutes to join the call and get themselves situated so that we can kick off with our panel discussion with Michael Ogilvie from Data Army and Gerry Stanley from Precisely. Great to have you here, fellas.


Gerry: G'day, all.


Michael: Thank you. Hello, how you doing?


Kylie: So if you're joining us, jump onto the chat. I know this is the favourite bit for proptechs around the country, but just jump on the chat. Let us know what proptech you're dialling in from or what part of the country. It's always fun to see how far and wide everyone's dialling in. We had a great RSVP level. Luke, founder of MicroBurbs. Luke Metcalfe. Luke, well done. Hey, Luke. Thanks for being first. Demonstrating audience participation – we love that. So if you're jumping on the call, just say g'day. Let us know. Christian from ArcSight in Newcastle. Hey, Christian. Brett from My Turn in Melbourne, but I haven't heard of My Turn before. Have we got them on our list? Great to meet you, Brett. Brett Peters. Okay, thank you. So if you're joining us, it's Kylie from Proptech Australia here. We had a really strong RSVP to this webinar, so we're just giving everyone a few minutes to jump onto the call, get Zoom restarted, and familiarise yourself with the tech. Jeff from Agent Best Practice in Sydney. Hey, Jeff, great to have you on the call. So if you're dialling in, please just let us know where you're dialling in from and what proptech. Owen from Applied AI. Awesome. Owen, you'll be able to possibly contribute to part of this conversation.


We might kick off with a bit of an Acknowledgement of Country. I'm coming to you today from the banks of the very beautiful Parramatta River in Sydney. And Proptech Australia acknowledges the Wangal clan, one of 29 tribes of the Eora Nation, and the traditional custodians of this land. Proptech Australia pays respect to elders past, present, and emerging, and extends this respect to all Aboriginal and First Nations people joining us on the call today. So we can still see people diving in, but let's kick off, 'cause we wanna make sure we get time to ask some questions at the end. Data solutions was one of the biggest categories in the recent Proptech Awards, and the number of people joining us on this call today does demonstrate how important it is as an issue for proptech. A strong and trusted data foundation is critical to successfully enabling AI initiatives across property, and through the awards and across the industry, we're seeing leading organisations already demonstrating tangible value through improved decision-making and operational efficiency. So while most proptech organisations recognise the importance of high-quality data, and many of you are out there creating structured data from unstructured data, many of us do still face a very common challenge: How do you transform that fragmented, inconsistent, or incomplete data into an AI-ready asset? So we're gonna be exploring this issue today and seeing some real-world case studies with the help of two experts in this area. So I'm joined by Gerry, Product Management Director at Precisely. Welcome, Gerry. G'day, great to have you here. Now, Gerry is a leader in location intelligence, data analytics, and transforming complex data into actionable insights. Gerry specialises in product strategy, customer engagement, and global collaboration to deliver precise geospatial solutions, and he helps businesses navigate the ever-evolving data landscape with confidence. So great to have you here, Gerry. Thank you.


Gerry: Thanks for having me.


Kylie: Awesome. And we're also joined by Michael, Director of Data Army. Data Army is a leading Australian data consultancy that helps organisations turn data into a strategic advantage. Michael has deep expertise across property, real estate, and proptech, and has delivered data-driven solutions that improve decision-making, drive innovation, and create competitive advantage. He is passionate about helping the industry harness modern data platforms, analytics, and AI to solve complex challenges and to shape the future of property. So welcome, Michael.


Michael: Hi, everyone. I see some familiar faces and 10 new names on the list as well.


Kylie: Oh, Matt Webster's joined the call. Hey, Matt. Great to see you here. Before we kick off, just to give everyone some context, Gerry, can you just give us a quick elevator pitch or summary of what Precisely does?


Gerry: Yeah. So we're basically all about trusted data for AI analytics and business outcomes. We're a global data integrity software company. Our technology stack is used by 95 of the Fortune 100, which probably gives you a bit of an idea of our size. And we help organisations trust, govern, enrich, and activate their data so that they can power AI analytics, operational efficiency, better business decisions, and that type of stuff with complete confidence.


Kylie: And Michael, what's your elevator pitch for Data Army?


Michael: Yep, we're not global. We're Australian-based, with offices here in Sydney and Melbourne, but we've got a bunch of consultants that work with organisations to help get value out of data. Usually that's consolidating it, building analytics and AI solutions over the top of it so ultimately things can change in the business that help the bottom line. That's what we do.


Kylie: Fantastic. So let's dive in. Most organisations know that they need high-quality data, so why is it such a challenge to achieve that in practice? Who wants to answer that?


Michael: Me – I'm happy to go first, 'cause we've worked with real estate data for quite some time now, and a lot of our clients are real estate-based. What's not commonly understood with real estate data is it's different to a lot of other data that you might be working with – for example, marketing or sales, or HR data. That's generally fairly predictable and there's quite a lot of alignment and generally a single system you might deal with. In real estate, we deal with hundreds, sometimes thousands of data sets. There's no conformity, there's no standards. Everybody can do almost what they feel like, so it makes it much more challenging. And even working within a single state in Australia, it can be like working across the globe in just the volumes of data and the diversity you deal with. So that's probably the biggest one, and there's no central knowledge base where absolutely everything's captured. Often these lessons are learnt the hard way, which is just getting in, rolling up your sleeves, and doing a huge amount of data analysis. I think the follow-on challenge from that is it is a large and difficult problem, and most businesses don't invest in it correctly. They don't put enough dollars behind it to get the outcome. You can use a lot of vendors – there's some vendors on this call, which is good – to help get you a lot of the way and take some shortcuts, and I highly recommend that. But ultimately there's a lot of work that still needs to get done to get good value out of data and often, like I said, not invested in heavily enough, I don't think.


Kylie: And look, I'm gonna confess that everybody having to invest the same amount of energy to bring it up to a standard to do something with feels like an issue that we should all be collaborating on solving, but let's just put a pin in that for a minute. So how can proptech organisations ensure that they've got a complete view of properties? And what are the attributes that are most relevant in addition to property details? 'Cause property details are expanding too, aren't they, at the moment?


Michael: Gerry, do you want me to handle that one?


Gerry: I can jump in onto that one. What I'd probably preface this with is most proptechs don't actually have too much of a lack of data to start with. They normally know a lot about the particular things they're currently looking at. It's really more so around how they can use it better, how they can understand it, how they can trust it. Once they actually get through that, once they can understand and have that data right, they start then looking at: What can I add to it? The big trends that I'm seeing are around movement profiles. So it's vehicle movement, human movement – those types of aspects of what's happening around the site, and how can you tell the difference between two different sites that look exactly the same from a property point of view? Which one is a better investment from, I suppose, the external factors coming in? The second one of those – again, external factors – is really around demographics. And it's not just what's the population around, what's the growth of population; it's what's the mix of population, the consumer styles, that type of side of it. And then the third part's probably then around business profiles. So rather than just the movement of the people, it's what are the businesses and what are the topological factors, I suppose, of what's in that region around it that changes it.


Kylie: So what you're saying is that we're seeing not just focus on the features of the property, but now also the people who live in the property and then the community that exists around the property, because that's impacting lifestyle and therefore the people and the features in the property, right?


Gerry: And that's what attracts the next set of people to a property. So if you really wanna know, if you wanna market to the areas that you wanna go to, that's where you've gotta start – what's around you, who do you wanna try to then bring into that location? What does that have to look like? And that's not just about what the property's gotta look like; it's accessibility to services, the additional services within the area. It's that whole component. And not very many people wanna live right on a busy road, but if it's a high-rise apartment, maybe you do wanna be on a busy road.


Kylie: Just tight so you can't hear anything, yeah. So one of the guys said... Oh, sorry. Michael, did you wanna add something to that?


Michael: Yeah, I was just gonna add one caveat in terms of attributes, which is my engineering background coming through, and also talk to that first point around standards. Integration IDs within your property universe are a key attribute to have. In order to grow the number of data points you can collect and store against a property, you need to be able to integrate it. So it's a really overlooked aspect of databases that makes it very difficult to do anything at scale if you don't have it.


Kylie: So would that be one of the common data issues that you guys are seeing in property or location-based data sets today that's missing from them?


Michael: Yeah, it's similar to that first point. Every originator of data often has a particular purpose for creating that data, and it's not always about integrations downstream and a broader view of the market. It's usually more singularly focused. So that means you need to take that data and clean it. Providers like Precisely across their data sets will have a common integration ID that allows you to bring all their data together. But if you wanna tap into the hundreds or even thousands of other data sets, you need common keys. So there's a bunch of techniques around – like geocoding's a really common one – that untaps the huge amounts of data sets that you can append on. But there's other integration keys in the industry that aren't a standard, but are used as a de facto standard. And putting them onto your data set does provide a lot of capabilities.


Kylie: Okay, I'm gonna quiz you on that in a minute. Gerry, sorry?


Gerry: Yeah, so I agree with that one completely, so I'm happy you're gonna come back to it, so I won't touch on those linkages at the moment because it's a passionate area for me. I think some of the other things really are around metadata. As we start going into AI, the lack of metadata against data that people hold is a really, really big, I suppose, restriction to actually extracting the most out of the data that they have. And the other one is really around consistency across areas. So quite often, proptechs will know a really deep amount about a small subset, but not that same depth across such a broader area as Australia or a state or even an LGA. So I think those two are really important in terms of the data issues that people face as they go to try to expand or extract really smart insights out of their data.


Kylie: And so how are those issues typically gonna impact analytics or reporting or AI outcomes if they're not in place?


Gerry: So metadata's a really important one. When you go past metadata into semantic layers and bits and pieces that I'm sure we'll probably get back onto at some point, if you don't know what the data is, AI will tell you with complete confidence something that's completely and totally wrong. It's the biggest issue with AI. It's the reason why there's a default message about any AI tool to say that AI can be wrong. But AI will quite often speak with complete confidence. That is where the problem comes in, in terms of metadata in particular, and a lack of semantic layers when you start linking these data sets together. I've got massive numbers of examples where it just links to the wrong things and does the wrong things, and you only have to Google issues with location and AI to start with, but really, that's the problem. It's around trust. Everything with AI comes down to trust, and without metadata, AI doesn't know how to use data properly.


Kylie: Right, okay. So what's the best practice for setting that up?


Gerry: Starting with a really good base. Adding in metadata where you don't have metadata, or improving how that is. The other one, of course, is bringing in baseline trust sources underneath that actually create part of that baseline to start with, and pin your data to it. Then you'll start being able to use the benefits of the metadata associated with a land parcel or a building or an address, and you actually get to use that as part of a trusted foundation that you can pass back through – more like how that type of data's used in fraud analysis – but using it to actually test your own quality of your data sets.


Kylie: Right. So just as a follow-up question to that then, Gerry: What's the best practice for evaluating an external data set? What criteria should you be using if you're looking for third-party enrichment data?


Gerry: It's massive. I've been around this stuff for many years, and we've got suppliers all over the world for this stuff. There's five Cs that I have, and all these different people have different things for it. So for me, it's around firstly starting with understanding your use case. If you don't understand your use case and what you wanna get out of it, then just backtrack and start back there. Once you understand the use case, you need to understand the coverage that you need for it, then what the currency needs to be. So what's the freshness of the data? Consistency then is quite often a hidden metric. So yes, you can have coverage, but if it's not consistent across that coverage, there are issues. There's always a data quality component that then comes around correctness. So you can have coverage, currency, and consistency, and it can still be completely wrong, or not accurate enough for your use case – so correctness comes in. And of course, when I come from a company that's all about trust, credibility is the last point that I have in my five Cs. You have to really understand where that comes from. And if you have a supplier and you're getting data from them in some way, shape, or form, if they can't answer things around where the data's come from and the quality process it's been through, then you've got to basically set up your own quality checks for what that credibility is, and it gets all the way down into legals. The amount of times I jump down into legals with suppliers and all of a sudden the red flags start going up. So they're my five.


Kylie: No, they're awesome. So can you just repeat them again?


Gerry: Yeah, so coverage, currency, consistency, correctness, and credibility. And then comfort. And so credibility really gives you the last one.


Kylie: That comfort, okay. So how are Precisely's data and geo-addressing products – how do they support proptech use cases and AI-empowered workflows, like specifically what you guys are doing?


Gerry: Yeah, so really we look at those sort of five Cs and see what we can do in terms of what the content becomes. And you can sort of then take away the content parts – so breadth of data types and coverage, and making sure of things for the accuracy and the like. You get into three parts that I think have been the focus area for Precisely over the last sort of eight odd years that I've been here. The first one of those is around accessibility. So it's everything from transactional access to direct downloads and access via cloud environments. Now, they're three incredibly different go-to-market models, but quite often the size of a proptech will mean that they wanna go down one path or another. Or it could be that even if it's a really big proptech, transactional access to high-value data assets could be really important as well. And there's been a big sway across between people that have wanted APIs and they've realised that for consistency, that's really problematic. So then they wanna download data, and then they've gotta have engineering teams managing the data. So then all of a sudden they wanna go to cloud environments, when you talk about places like Snowflake and Databricks and the like that people use for that. The second one is then really around usability. And the usability is probably the most important, really. That used to just be data formats and projections and stuff in my spatial world. That jumped to linkages – Michael was talking about linkages before – the trusted IDs that you put in the data, and not just your own, but third-party ones that these things can link to. And the most important in the AI era really is around metadata and semantic registries that have to be there for that usability to actually shine. And my third one of them is data pipelines. So you actually have to get rid of the whole thought process that people just wanna access a piece of data and just access that data. It's actually part of the pipeline. So it's going from an address to a geocoder, to a property, to demographics, to risk, to movement profiles, to business information, but all in a single flow. And one thing that AI is pushing us down the path of is the fact that all this stuff happens very quickly and happens in a really high-quality way so that we can make informed decisions. Under the covers, if you don't have that usability component there, there is no way it's known that's happening.


Kylie: And then it's where AI goes all wrong. So, Michael, question for you – because I'm sure this is in your wheelhouse – onboarding third-party data can be pretty time-consuming, especially if it's not in good order. How can organisations accelerate that process and more easily maintain data relationships as things change over time?


Michael: Yeah. This is a topic I'm particularly passionate about. Coming from an engineering background, I've worked with pulling data in from all different places. I think it really depends on the use case, as Gerry said. If you need one record at a time, something like an API makes complete sense. If you need to download 15 million spatial records, an API is not usually an efficient way of doing it. So I think people need to make sure they're using the right approach for the right job, and making sure they're picking vendors that have multiple interface points. There's a lot of challenges in getting a project live and returning value. Moving data from one point to another is not one of those key value items. It's a necessity – you have to do it, but it should be as simple as possible. So marketplaces are a great one. APIs can still have their place, and so do MCP servers, but it needs to be fit for the use case. So that would be my advice there.


Kylie: Yeah. So can you share an example of a data enrichment project and how that's unlocked better outcomes for someone?


Michael: Yeah. So we won't name them – we haven't got permission to do so – but we worked with a large ASX business that deals heavily with property portfolios. They understood their properties reasonably well, but that's not the entire market. So they wanted to build out a complete market view. We used Precisely data in that instance to get a view of every single property in Australia, and then we enriched it with some more Precisely data, plus a whole wealth of other points of data. So they now have a deeper understanding of their own properties. They have an understanding of their competitors' properties. They have an understanding of properties they might want to purchase, and it's far beyond what they could have done with just internal data assets. And this is not just the case for an ASX business – proptechs have this same issue. They're usually up and coming. They have a subset. They don't have the complete picture of every property out there and to enough detail. So that's where we sort of look at completeness of the picture: Build a foundation that we can extend and have a pattern where... new data sets are coming through every day. You need a quick way of onboarding them and enriching your data.


Kylie: Fantastic. Okay. So moving on from foundations and what needs to be in place, the next step, I guess, is people wanting to execute and deliver – and that's where you guys sort of do a lot of work. Once you've got access to high-quality data, what does it actually take to turn that into a working solution, Michael?


Michael: Yep. So the cleaning part is the hard part in good-quality data. So again, you leverage trusted vendors as much as possible to cut down that task, but there's inevitably still a big component of projects that is that. But the next piece is integration, which we touched on. And again, whichever vendors are gonna reduce that problem for you with common identifiers, you 100% grab them as much as possible. There's a huge value in doing that. But after the integration's been done and everything sort of can talk to each other – the data sets – the next bit is we model the data. This isn't new. It's not something specific to real estate, but building canonical models that represent reality – so not a source system, not a real estate CRM or something like that – it matches the real world in terms of processes that exist. And then once we have that data all modelled in there, the sources no longer matter to us. We've got this canonical model. It's then building out – in the case of natural language interfaces, we might build a semantic model over the top. If it's ML, we'll start building further features to go feed those models. If it's document processing and that sort of thing, then we might look at how do we do efficient pipelines that minimise the tokens that we're using. But it all comes off that foundation of modelled data. But this is the step most organisations skip – they go straight to the AI/ML piece, and usually it's on a shaky basis. So it all...


Kylie: I'm gonna ask you a question without notice, but I know that you're gonna know the answer to this. So you mentioned two types of models, canonical and systemic – what was the second one?


Michael: No, no, no, I was saying canonical models in terms of modelling towards reality or building... like, when you pull the data in from a system, usually you'll keep it in its raw form, so it'll match that source system, whatever it might be. If you build your solutions off that, you're very tightly coupled to a system, and most businesses will change systems periodically, and your whole solution's been coupled. It makes a very big migration job. So we try to disconnect from those source systems, and we focus on the business, 'cause businesses do change, but generally sales happen in a similar way, accounts receivable happens in a similar way. All those sort of things are pretty persistent, so we model those things out. And then when someone from the business is asking a question on the data, or you need to build a model of the data, generally you wanna model an answer based on that business process, not some source system. So it's a switchover point where we stopped caring about the source systems and care about the business.


Kylie: So this truly is about setting the data free so that it's actually independent of where it's come from, and that it can go into any new model, I guess.


Michael: Yeah. And it's also what makes it extendable. Because new processes – that's fine – but if you acquire a business – this happens quite a bit in the proptech space – if an acquisition happens, you don't really have to fundamentally change your business models, like that layer of your data warehouse. You just map in more data into it.


Kylie: This is how we scale. Okay. So how does Data Army typically approach transforming fragmented data or sub-par data into something that's ready? Is that aligned to what you've just talked about, or is there more to it?


Michael: No, that's a big part. So the precursor to building those canonical models is usually understanding the business. It's extremely easy to get stuck in a project that goes for years and not deliver any value, because you're trying to solve every problem under the sun.


Kylie: And real estate – well, agents, who are our clients, often send us down that path, right? That's very common.


Michael: Yeah. It's very easy. So you have to focus on business problems and what it takes to solve them, what their risk profiles are, what can they accept, what puts the business in a better position than they are today, and then go backwards and start building out the models and sourcing the data to populate that. So that gives you a very outcome-focused view of it and gives direction. I tend to find if you just start by bringing data in and start modelling, start doing all these things, you get lost in the weeds and nothing comes out of it, and projects tend to stall in that situation. So yeah, we start with the business, work our way back. Generally, find the easiest possible way to get data in – bring it into a central store with the least amount of friction possible – and then we go down that modelling task and sort of iteratively cycle between cleaning the data, integrating the data, modelling it, unlocking value each time.


Kylie: Yeah. So have you got a case study that you can talk us through on how you did that?


Michael: So we can name some. I think we might have mentioned previously, we've had a long relationship working with Archistar as a proptech. Now, they do work with thousands of data sets, both here, the States, Canada, UAE – lots of different places. So again, it's a very simple one – like, we could be looking at geospatial data sets for the next 10 years, no problem, but that's not gonna get them anywhere as a business. So working alongside their team was really taking that approach of: What feature or functionality in their platform are we trying to unlock, what's the data involved, and what are the error margins that are acceptable to get a business gain? So, yeah, we've worked on that for a long time, and we just keep iterating. And we take selections of whether we wanna increase quality some months, or unlock new functionality, find new markets. They're all variations that we sort of do, and it's a balancing act. But it's very fine-grained in terms of how much control on the task we do. We never get into a situation where we don't know what's happening. We always know this amount of effort here is gonna result in this. So, in that position, I think most businesses wanna get in – is a certain amount of investment yields a particular return.


Kylie: Yeah. Okay. And look, Archistar's a great example, because they scaled very quickly into external markets, I'm guessing now with your help, right? Because of bringing in all of those data sets separately. So were there any unexpected benefits that came from improving the data foundation around that stuff? I mean, you've just talked about knowing what the outcome's gonna be and how much you're gonna invest in it, but are there any other happy surprises that come out of it?


Michael: Two main. One I wouldn't say is unexpected – it was always a goal, which is a lot of tasks within development have to be done, like a lot of these cleansing tasks and things like that, data acquisition. But they aren't the outcome that the business is after; it's a means to an end. So automating that meant a lot of the resources within the Archistar team got freed up to do much higher-value items that ultimately do deliver value for their customers. So I wouldn't say that's unexpected. It was expected, but it became a reality.


Kylie: Was the impact bigger than you thought it was gonna be once it was revealed?


Michael: Yes, I think so, in that... like, being able to tackle new markets. Archistar is still in a growth phase. They're not a billion-dollar organisation with 1,000 headcount. So to tackle that many markets with the size of business that they are, it's impressive. You need to have scalable systems to do that, and resources need to be able to still deliver new functionality which the investors expect and also their customers expect, while still keeping a lid on all this data that's constantly in motion. So it is, I guess, unexpected in terms of being able to still run at that scale for the size of business, so it's impressive.


Kylie: So what I'm hearing as part of this conversation – and Gerry, I've got some questions for you coming up. But AI sitting at the end of this process, right, that we're talking about now, about getting your data right, getting it structured properly, getting it ready for scale. But what are some of the practical AI use cases you're seeing – that both of you are seeing?


Gerry: So I can probably kick off on that, because your statement that AI is coming at the end of these things... wouldn't have been too long ago and I probably would've agreed with you. But AI is not coming at the end anymore. AI is across so many touchpoints, and that's been progressing over time, of course. But everything from your situation use case analysis right at the beginning, before you even go looking for data, is AI-driven. Data discovery as to what's within data, creating metadata out of it – all AI. Pipeline creation now can do a whole lot through AI. MCP servers are a perfect example as to stitching those sorts of things together. Governance models, all AI-driven. Yes, they probably need a lot more individual indication and analysis at each point, but I don't know any top-tier governance model that isn't actually being powered by AI to do it fast and to get it across whole organisations. And it continues right through quality verification, doing enrichment. And then you finally get to the bit where most people think that the AI is, and that's in the analysis and the predictions and the dashboards and the reports and everything else. There's a whole heap of AI across every part of all of that. I will always come back to: The only way you actually get through all of that is by building metadata during it, building the data linkages, getting data standardisation, advancing into semantic registries – which most people don't wanna touch, but it's actually where the gold actually comes out a lot better. We're going all the way through to pre-built narratives across our data. And the reason being is our customers, as they're doing this stuff, the more we do in that background, the less hallucinations they're seeing. And we all like talking about agentic AI. You go to any conference and it's agentic AI, agentic AI. Just try using slightly different wording and see if you get the same outcome. If you're not doing that, that's not agentic AI. So that AI enablement, yes, used to be at the end. The end bit is the fancy, glitzy part of AI – nice shiny things. But the hard work is under the covers. And AI is driving that at an alarmingly fast rate.


Kylie: Okay. No, I will... thank you. I will consider myself appropriately schooled. And I hadn't actually, of course... like when you were saying that I was like, "Of course." But so how do you get the AI right at the start while you're cleaning the data? Like I feel like we just went into an eternity loop there. I was gonna... How do we pull ourselves out of that?


Michael: Touch on that, which is... so there's a few topics that we discussed today that tie in together, which is: If you're picking a vendor that's providing data with metadata, then AI can help further. If you've got data sets without metadata, and it's unreliable and you don't necessarily know when it was created and it's not always accurate, using AI in the development process or any of these stages generally doesn't go too well. Or it can be an assistant and it can help, but you gotta assume that it's gonna be wrong, the same way a human would be wrong. If you pointed a human and said, "Here, have a read of this," they read it, came up with an analysis, and said, "Oh, by the way, I forgot to tell you that data's 10 years old" – it's all relevant information. So it's gotta be used cautiously. It does add value and is speeding things up, but it's not quite the same as software development where often you're working with a blank slate and you can create something, and as long as the interactions are correct, there's a known output. Data often – or sorry, the variable output is okay, but in data generally there's like one output and you don't necessarily know how to get there. So being creative in its outcome is not always a good thing, particularly in finance.


Kylie: And it can... Okay, so I'm also gonna... So can you use AI to support creating the metadata if you've got a data set without metadata? Can it assist with that?


Gerry: Yeah, definitely. There are metadata development tools that are based on AI principles. Doesn't mean you can trust everything that comes out of them. It gives you a starting point. It gives you an initial structure. That's where third-party data sets become really important, because the only way to really test that, if you don't have another trusted source to put against it internally, is you actually wanna start pinning it against things. You wanna question it. It's almost like putting an AI tool into defensive mode and saying, "Don't tell me I'm good or that it's a great idea anymore. I want you to tell me why it's a terrible idea." It's reversing the use of AI from something that we think is just gonna create something new and wonderful for us, and instead using it to actually query and to try to tell you why something's wrong. And guard against the risks. Best way to use AI ever is to have it so that it's defensive, not proactive.


Kylie: Okay, awesome. So for organisations that know that their data isn't where it needs to be, where should they start?


Gerry: Go to Mike is what I would say. Contact Data Army. Start at Michael. Michael first. That's where I'll start.


Michael: Yeah, no problem. I think sometimes it's fundamentals that most organisations skip. And when we go into a lot of organisations, that's what we see. So things like putting your data in one spot. It's really common that it's spread all over the place. We don't have access to this, we don't have access to that. Most of the time, there could be some legitimate reasons governance-wise why data can't be shared. A lot of the time it isn't. So the concept of a data warehouse where all your data comes into one spot – it's been around for decades. It's still true. And, I mean, even Anthropic published some papers – they still use a data warehouse. A lot of the cutting-edge organisations that most people think have done away with all this old tech still use it, because it works and it makes the whole surface plan of what you're dealing with significantly easier. If you've got the data in one spot, you can integrate in one spot, you can clean it with this common set of tools. So that's one of the first ones. And then the next one I'll jump to is really understanding the business and then the data that marries up to the business, 'cause again, you can go off creating things. If you can't instruct the AI, there's no way the AI is gonna be able to come up with a valid answer. It's almost impossible, other than it just guessing and having to guess right sometimes.


Kylie: And so how can businesses assess if their current data foundation's fit for purpose?


Michael: I split it into two areas. One is: Is your platform correct? Because sometimes we come across organisations and they're using their data infrastructure or their technology infrastructure, and it's just difficult to roll out AI. So you've got all this friction happening within moving data or getting services to connect. And the bigger the organisation, actually the bigger the problem this is. So AI should be pretty seamless. You should be able to extend it into your current way of working without a huge change or shift in what you're doing. So I mean, we work with all the modern cloud data warehouses. Tend to preference Snowflake. It works for us a little bit better. But in those sort of environments, they can handle massive amounts of data and can do geospatial things. They can do AI. All these different concepts, and a lot of the providers work in this platform space now rather than a database. So if organisations are still on a database and then trying to cobble things together, it's a more difficult way of operating. Not impossible, but more difficult. And then in terms of your solution, we look for patterns – lots of patterns that allow you to scale. Because you can build a prototype, it can look cool, it can work pretty well, but then you start adding more and eventually it gets into a grind. Building it out takes far longer than it should. Often an upfront investment into building out some patterns means you can add a lot more functionality when the business demands it.


Kylie: And Gerry, have you got anything to add to that?


Gerry: I think one of the other things is also bringing the human side of me back into the loop – test things that you know, but don't build the models off everything that you know. So you need some test stuff separately. And when I talk about testing things that you know: When Google Maps first came out, what was the first address that people put into Google Maps? Their home address. Always is, because it's what they know. It was the first one that they typed in. It was the first one they zoomed in onto. Then they probably started looking at the schools. Then they started looking at the shopping centre. It's all around things that you know. This type of data's no different. So test against things that you know, but don't test all of your data that you know in the modelling component. You've gotta separate out some separate stuff that you also know is right, but use that as testing outcomes as to what it is. Don't throw everything just at the beginning.


Kylie: Can you give me an example of that?


Gerry: So one of the coolest activities I got to do was creating all the buildings across Australia with my prior employer, and you have to set up a test data set. What does a building look like? So all of a sudden you end up with a test data set: This is what buildings look like, this is what the pixel models for machine learning need to run like to train the model. But once it's through the training of the model, you've got to test the model. Yes, we do visual tests with large groups of people, but you also want machine-based tests. The best way to do that is having a set of data that you also know is perfectly accurate, but you actually run that against what the outputs are that are coming out of these processes. It's the same thing. The challenge, I think, when you get into proptech is the pace at which they're trying to move. The very intense knowledge they know about the smaller activities that they're doing – that expansion, that's an absolute killer. So anything that you're doing for assumptions or models on one side, just don't throw everything into creating those models. That's what I see a lot of people do – they go, "I've got all of this, I'll piece it all together," but then they've got nothing to test it at the end if the model's already built on it.


Kylie: Right, okay. So, and look, if anyone's got any questions for Michael or Gerry, please pop them into the chat and we'll get to them very shortly. But how can businesses assess whether their current data foundation is fit for purpose? They're listening to this podcast, listening to this webinar and going, "Oh God, oh God, oh God." No?


Gerry: So my first thing is there's lots of AI tools out there that get people moving very, very fast. Treat everything as though it's incorrect until you can prove it's correct – that's probably my starting point. There's lots of data. Even when we've got clients like MA Financial – one of the largest of their kind in Australia; there's about $180 billion, I think, in their managed portfolio somewhere – they're massive. Where did they start in this process? They got one small part of their business, they came up with the use cases, they got that data in. Testing that data, doing what the enrichment is against it, looking to see how that actually flows through into the rest. It's not about eating the whole elephant in a single bite; it's about picking those first ones, getting that part of the modelling right, then going to the next one. And that doesn't really matter whether you're a large business or a small business. You've got to do it in that moderated step pattern, 'cause like Michael said before, otherwise you'll have this whole lot of blur and you get caught up trying to fix each individual element that breaks another element, that shows you that another element's wrong. Pick a core use case, focus down onto that, know what you're expecting to get out of it. AI is incredible. It truly is changing people's lives. But it also lies, and it lies with confidence. You have to make sure that there's a human element going, "Does that look right?" Without that human element, you can't trust what's out the back end of AI, unfortunately.


Michael: I can extend on that human element part as well, which is: Building these use cases, it's getting easier and easier, and you can go to release them quite quickly these days. The harder thing is actually operating them. And operating in that – it's the same as a SaaS application that you build. The more people use it, the more issues surface. More eyeballs, more questions, diversity – things will come up, and you need to be able to iterate pretty quickly. So what are the mechanisms you have to track user queries, to test for drift in the data, all these things that are gonna derail the solution? Because you can build trust quite quickly, but you can also lose it even quicker, and then trying to get people to reuse it again... So big components are: How do you collect this metadata or this context about the data? Who owns it? How do they maintain it? What mechanisms are there to make sure it stays up to date as well? It's not a difficult task to do it, but pinning ownership on people around data has always been a problem, even before AI. So getting somebody to own a data point indefinitely is a tricky one. So there's a lot of internal operations or ways of working that need to get established in organisations to make sure, one, you can get it launched, but then keep it successful.


Kylie: Yeah, okay. And so, Michael, what are some... You mentioned Archistar before, but what are some more practical use cases that you've seen, or some of the best ones that you've seen in the property and related industries?


Michael: There's still a few that are more common than others. So AI is a broad field, and there's ML, and ML's been around for over a decade, so that continues to deliver value. I think most people think of AI as the LLMs and what they're doing. So, from before AI days, still getting data into the hands of an exec or managers has been difficult. Not in terms of a technology, but comprehending, you know, this 30-page report – what's it actually really telling me? What's the bottom line? And maybe they are able to read that 30-page report, but do they have the time to do it? So summarising, augmenting traditional BI, I'll say. Not replacing and not summarising, but augmenting traditional business intelligence with an AI overlay is a big area we see across all of our customer base. So, great, I've got a 30-page report – maybe that can be 10 pages now, and can you give me a couple of summaries that I can digest quite quickly that link through? And then they go into a normal business intelligence flow from there. So that's a big one. The other side is unstructured data processing. It's coming through definitely in the last 12 months, so PDFs, images – all those sort of things. It's growing. I think there's a lot more there untapped, but more and more organisations are starting to do that, and the big challenge there is doing it at scale. It can be costly sometimes. So if you're processing millions and millions of documents, how do you do that efficiently and still reliably get an answer? And then we sort of touched on the development process. That's probably had faster adoption, which is people using copilots to help build these data platforms and solutions. That's torn through the industry quite quickly, not just in real estate, but throughout all development.


Kylie: Are they using Copilot or are they using Claude? Did they even use Copilot, really? Sorry.


Michael: Sorry, Copilot in a generic term of an AI assistant to help you with development, so yeah.


Kylie: Any one of them. The future of Microsoft is well known. Sorry, sorry, Gerry, interrupting.


Gerry: Any one of half a dozen of them. Depends on what you're wanting to do, which one you choose quite often.


Kylie: Fair enough. We know Claude was very busy for the Proptech Awards. Very, very busy. So look, have we got any questions? Or I know that you wanted to just share a screen, Michael. Did you wanna...?


Michael: It was just more if people wanted to follow up with us. I will quickly share. We've got a QR code or a GoTo form if I can work out how to get my computer to work.


Gerry: Let's do that. And whilst Michael's playing with that, there's a Snowflake World Tour in Sydney next month. If people are interested in what Michael's been saying about Snowflake – and myself, I suppose – if you can get a gig to go there, it is quite a cool get-together for that as well. And either Michael or myself can get you more details if you reach out.


Kylie: Okay. So we can reach out to you if we're interested to find out more about that. And we'll share out a link to that when we send out the replay to everybody too. That would be great. If people are thinking on the call about how to improve their data foundation or enabler, what is the next best step?


Michael: Oh, sorry, go Gerry. I'll go after you.


Gerry: I'll put one thing out there, particularly for proptech. One of the common issues I see is what a property is defined as by different companies. Some will pin it to an address, some will pin it to a land parcel, some will pin it to a building, some will pin it to a particular construction of an item. It's constantly different, and Michael spoke a lot about those data linkages. You've gotta work out how you're gonna link that data together so there's a single thread across anything that you're doing in a property area. I actually call it an intelligence layer myself. There is some sample data that looks for that type of stuff off the Precisely website, but there's also samples in Snowflake. If you don't know what that sort of linkage looks like, understand that linkage, understand what that looks like across your total portfolio so that you're singing from a nice strong base to start with. All the rest of the enrichment on top is great, but you've gotta actually define the way in which you're holding that data and referring to that data first.


Kylie: Yeah, I guess, look, this follows the same foundation that we have in property, right? You can't paint a wall or wallpaper a wall if you haven't got good foundations. It's all just gonna pile in on itself afterwards. So getting those foundations right. And I guess what I'm hearing today from the two of you is that getting the structure right, getting the metadata right, getting the integration layer right so that the data is known, can be trusted, can plug into other data sets easily, and transparently so that you can see what you are or are not dealing with – that's absolutely really key. But you don't have to try to solve that problem inside your own head. What I'm hearing is that by bringing in some of these external data sets, you can actually use them as the model, and making sure you trust them obviously first. But you can then use them as the model that you can then start to adapt to your own data. Would that be a fair assessment?


Gerry: Yeah. Correct. So in Precisely we use address as the centre of the universe. And some of our customers use property boundaries; others use cadastral boundaries. When you get into banking and the likes, that can get to volume folio because it links to loan documents better. It's actually just understanding what those different reliances are and how you can pin your data together most effectively by knowing them.


Kylie: Yeah. Awesome. Okay. Has anyone on the call got any questions for Michael or for Gerry? If you have, pop them into the chat, we'd love to... We've just got a few minutes before we wrap up. Just while we're waiting to see if anyone pops any questions in, Michael, what does a data integrity assessment involve?


Michael: Yeah. So we often, as part of our services, do an advisory piece where we often will go in, look at current platforms from a technology point of view, from a solution point of view. We generally will meet with the data teams, the stakeholders they work with, try to understand the pain points. We look at patterns that have worked and have not worked in some of our clients, where things stagnate, and then we'll give generally a set of recommendations in terms of a report. That often happens in a little under a week, generally. Depends on the size of the organisation and how complicated it is as to how much time. But organisations are each unique, but not that unique. Humans have a lot of commonalities in how they interact and cause friction points, and data has common friction points. So generally you can just spot patterns, and then map a solution that has worked previously.


Kylie: Cool. Okay. Awesome. Well, look, just gonna double check whether we've had any questions. Anyone? No, I can't see anything that's come up. But look, just to wrap up, is there anything that you wanna add that we've not covered off on yet?


Michael: No, I would just like to emphasise the importance of truly understanding your business and truly understanding how the data maps to it, because it is a fundamental step for AI enablement to be able to explain it to the AI. And AI can help to build that, but it generally doesn't get it right. No one knows your business better than yourself, so you need to be able to articulate it very clearly and keep that understanding up to date. It's alluring to jump straight to something like Claude, plug it in, you get a result, and it feels good that it's so quick, but you've skipped a bunch of steps, and ultimately it'll bite you somewhere.


Kylie: Yeah, that's a fake dopamine hit, everybody. Don't fall for that. Don't fall for the tick of the check.


Gerry: And I liked telling you that it was a great idea, and that was very perceptive of me. I like those positive reinforcements.


Kylie: It is, it's very beguiling, isn't it? I think too, this is actually really reassuring, right? Because it is actually understanding the real problem that we're trying to solve for. That's the human creativity that has the deep contextual knowledge that hasn't necessarily been articulated properly in a structured process yet, which is where the humans absolutely stay core and key to the ecosystem.


Michael: Correct. And the funny thing is – and again, this is not a new concept – humans have always had a difficulty recounting the knowledge that they have. You can recall a certain amount of it, but there's a lot that you know that you don't know that you know. Or that you can't remember that you knew. And there's even techniques not only using these AI tools as a copilot, in that you can use it to help pull out this knowledge that you have. But again, vibe coding's a big thing in the development world. If you're trying vibe coding and you're not having those interactions, you're probably conveying half of the context that you truly know.


Kylie: Yeah. Awesome. Well, look, Gerry and Michael, thank you so much. This has been an absolutely fantastic conversation. I'm gonna hit you both up to be involved in the Proptech Forum on November 13th. Please everyone pencil that into your diaries. It's a Friday. So Tim Castle's basically saying what is a property is one that he's had to deal with a lot too. Universal problem. What is a property? And it's... Even I remember from CoreLogic in Melbourne, when you had a dual property called something in Melbourne, which was called something completely different in every other state of the country. And so I remember, 15 years ago we had a whole bunch of properties that we hadn't counted 'cause they were not recognised by the system.


Gerry: Easily done. And how do you connect those then to what the rules and restrictions are underneath each one of those as well. And it all comes back to understanding what that complete little micro system is between address, parcels, properties, ownership. Those lines are all quite blurred at times, and incredibly fragmented across every jurisdiction.


Kylie: Yes, 'cause it's our federated model, and it's awesome. It keeps us all very busy. So look, guys, thank you so much for your time. It has been a fantastic conversation, and looking forward to diving into it again as part of the Proptech Forum. I think too, this is also a great opportunity to mention while everyone's just wrapping up on the call, the energy efficiency data set that Proptech Australia is behind for the new data coming through for energy efficiency. There's a data standard around that – we are trying to come up with a common standard for the new energy efficiency features that are coming out. So it is free for Proptech Australia members to sign up to that data set. Please do; it will give you a common architecture around what things are called and how to describe them and how to use them. We'll send out a link to that in the show notes as well. Thanks, Michael, thanks Gerry, for a fantastic conversation. And thank you to Data Army for making this possible. We will see you soon, very soon on the next proptech panel. Thanks, guys. All right. Thanks, all.


Michael: Thank you. Thanks, Gerry.



bottom of page